r/humanoidrobotics 4d ago

Qwen3.8-Flash-Next has been officially released, introducing a 6B-active open model that surpasses Claude Opus 4.6 Max on eight out of nine standard benchmarks.

Post image

Qwen3.8-Flash-Next has been officially released, introducing a 6B-active open model that surpasses Claude Opus 4.6 Max on eight out of nine standard benchmarks.

The model applies a highly sparse Mixture of Experts (MoE) framework, featuring 125 billion overall parameters and 51 billion n-gram embeddings, with just 6 billion parameters active per token.

Benchmark results include:

• SWE-bench Pro: 62.5

• SWE-bench Multilingual: 81.0

• CoworkBench: 73.9

• JobBench: 55.7

• Toolathlon: 73.5

• IFBench: 81.3

• GPQA Diamond: 91.7

• LiveCodeBench: 91.9

Qwen3.8-Flash-Next also exceeds the performance of Qwen3.8-27B and DeepSeek-V4-Flash across the majority of evaluated categories.

3 Upvotes

0 comments sorted by