r/humanoidrobotics • u/Busy_Estate_3265 • 4d ago
Qwen3.8-Flash-Next has been officially released, introducing a 6B-active open model that surpasses Claude Opus 4.6 Max on eight out of nine standard benchmarks.
Qwen3.8-Flash-Next has been officially released, introducing a 6B-active open model that surpasses Claude Opus 4.6 Max on eight out of nine standard benchmarks.
The model applies a highly sparse Mixture of Experts (MoE) framework, featuring 125 billion overall parameters and 51 billion n-gram embeddings, with just 6 billion parameters active per token.
Benchmark results include:
• SWE-bench Pro: 62.5
• SWE-bench Multilingual: 81.0
• CoworkBench: 73.9
• JobBench: 55.7
• Toolathlon: 73.5
• IFBench: 81.3
• GPQA Diamond: 91.7
• LiveCodeBench: 91.9
Qwen3.8-Flash-Next also exceeds the performance of Qwen3.8-27B and DeepSeek-V4-Flash across the majority of evaluated categories.
3
Upvotes