Alibaba Qwen Open-Sources Qwen3.8-Flash-Next Architecture
Summary
Alibaba’s Qwen team has released the Qwen3.8-Flash multimodal MoE model and open-sourced the Qwen3.8-Flash-Next architecture, an early foundation for Qwen4. The 125B-parameter model activates 6B parameters per token, supports 260,000 tokens natively, and can extend to 1 million tokens with YaRN. Qwen says training costs are one-ninth those of Qwen3.7-Plus. Its architecture combines GDN and QSA attention, gated residual streams, N-gram Embedding, and mixed Muon-AdamW optimization. Weights and a technical report are available on Hugging Face and ModelScope.