KIT_IWSLT26_IF_ LONG_ unconstrained_ primary
We participate in the unconstrained data condition. Our system is based on an end-to-end multimodal model built on Qwen2.5-Omni-7B, trained jointly across all target languages.
We employ task-balanced data interleaving with temperature-based sampling to balance heterogeneous task distributions and mitigate dominance of high-resource tasks, including ASR, speech translation (ST), spoken question answering (SQA), speech summarization (SSUM), audio captioning (ACHAP), multiple-choice QA (MC), and general instruction-following tasks.
At inference time, we apply an additional re-ranking step over n-best hypotheses to improve final prediction quality. This re-ranking is applied only for English and Chinese.
| From\To | de | en | it | zh |
|---|---|---|---|---|
| en | N/A | N/A | N/A | N/A |