An arXiv study finds that post-training ternarization can reduce Qwen3-4B storage substantially, but capability and perplexity worsen, while faster inference remains unproven.
AI News
The latest AI releases, research, products, and industry updates.
The latest AI releases, research, products, and industry updates.
With your permission, we use Google Analytics to understand page visits and whether key workflows succeed. We do not send prompts, files, account identifiers, model identifiers, balances, costs, or raw errors. Privacy Policy
An arXiv study finds that post-training ternarization can reduce Qwen3-4B storage substantially, but capability and perplexity worsen, while faster inference remains unproven.
Hugging Face’s page highlights Qwen/Qwen3.8-2.4T-A95B, a 2.4T-parameter text-generation model. The supplied text reports a recent update and approximately 21.9K uses and 1.18K likes, but provides no benchmark or comparative findings.
Alibaba’s Qwen team has released the Qwen3.8-Flash multimodal MoE model and open-sourced the Qwen3.8-Flash-Next architecture, an early foundation for Qwen4. The 125B-parameter model activates 6B parameters per token, supports 260,000 tokens natively, and can extend to 1 million tokens with YaRN. Qwen says training costs are one-ninth those of Qwen3.7-Plus. Its architecture combines GDN and QSA attention, gated residual streams, N-gram Embedding, and mixed Muon-AdamW optimization. Weights and a technical report are available on Hugging Face and ModelScope.
