Back to News
User submissionnews.aibase.com

Qwen Launches Qwen3.8-Omni-Flash with 1M Context and 26% Average Benchmark Gain

Summary

Qwen launched Qwen3.8-Omni-Flash, a new native multimodal model that accepts text, image, audio, and video inputs and supports a 1M-token context window. It is available for testing on the Qwen AI platform. The model extends Qwen’s earlier capabilities in coding, text-based knowledge work, and GUI operation, while placing greater emphasis on agentic workflows centered on audio and video, including video editing, music-video creation, film and television production, narration, multimodal summarization, and audio-video dialogue. Across 30 evaluations, it achieved an average score more than 26% higher than the previous Qwen3.5-Omni-Plus. Reported gains included 36.5 points on WildClawBench-MM, 22.3 points on AgenticVBench, and a score of 69.6 on UniClawBench. Other improvements covered long-audio understanding, video understanding, audio-video captioning, and meeting transcription: LongAudioSpan rose by 8.3 points, OmniVideoBench by 9.6 points, and OmniCap-IF CSR/ISR by 8.5 and 14.1 points. On AliMeeting, DER and cpWER fell from 88.11/89.61 to 3.35/17.18. Qwen says its audio-video capabilities are close to Gemini3.8Flash and that its overall audio capability exceeds it. API audio-input prices fell by more than 98%, while audio-video input prices fell by more than 93%. To support long-running workflows and real-time interaction, Qwen expanded Qwen-MM-Plugins and open-sourced Qwen-Live Harness. In Agentic Understanding mode, OmniVideoBench accuracy increased from 63.4 to 67.8 while token use dropped from 145,736 to 79,117, a reduction of about 45.7%. The team also reported an autonomous research-development cycle that selected evaluation sets, built data, and completed four iterations within 12 hours, producing 3,413 training examples and reducing Sichuan-dialect character error rate for Qwen2.5-Omni-3B from 25.79% to 15.30%. A Realtime version was released at the same time and is described as the first native multimodal model with sound-source localization.