Qwen Releases Qwen3.8-LiveTranslate with 2.3-Second Simultaneous Translation Latency
Summary
Qwen has released Qwen3.8-LiveTranslate, a new simultaneous-translation model upgraded across translation quality, latency, speaker identification, and speech synthesis. Its audio-text interleaved architecture combines streaming understanding, text output, and speech generation in a single causal sequence, reducing per-character latency from 2.8 seconds in the previous generation to 2.3 seconds. The model adds real-time speaker separation for conversations with multiple alternating speakers and can preserve each speaker’s vocal timbre in translated speech. It also outputs the original and translated text in the same frame for synchronized bilingual viewing, while long-context disambiguation uses prior conversational context to reduce ambiguity involving proper names and references. A Hybrid MoE design separates the Thinker module, which handles understanding and translation, from the Talker module, which synthesizes translated speech while retaining the original voice characteristics. On the 14-language-pair, multi-speaker, long-audio Omnilingua-MSpeaker evaluation set, Qwen reports better translation faithfulness, fluency, conciseness, and speaker-separation error rates than mainstream real-time simultaneous-translation systems. The model currently supports 60 languages, and its API is available on the Qwen AI platform. The team says future work will target lower end-to-end latency, cross-session long-term memory, and broader language coverage.