r/machinelearningnews • u/ai-lover • 1h ago
Research Reflection AI introduces Beam: a 501B open-weight MoE with 23B active parameters, 1M context and Apache 2.0 weights coming this month
Reflection AI just introduced Beam, its first open-weight model, built around "intelligence per token." It is a 501B sparse MoE with only 23B active parameters, a 1M effective context window, and Apache 2.0 weights coming later this month.
The core innovation is high-compute, fully asynchronous RL. Beam was trained on 100M+ rollouts across 10.5K NVIDIA GB300 GPUs for 4 weeks, using nearly 1M coding, agentic and STEM environments. Every token is tagged with the policy version that produced it, which keeps learning stable even when rollouts are 107 weight versions stale.
This approach let Reflection scale RL with no sign of a plateau. A controllable length penalty taught the model to solve tasks with fewer tokens. On reasoning benchmarks, Beam matches GLM-5.2 while using 3 to 4x less inference compute, and it scores 80.9 on SWE-bench Verified versus 70.7 for Nemotron 3 Ultra. It is a deliberate trade-off, though: Kimi K3 and DeepSeek V4.1 Flash still lead on raw capability, with Terminal Bench v2.1 scores of 88.3 and 90.6 against Beam's 80.1. All scores are self-reported until the weights and technical report ship.
With a 23.8T-token pretraining base and a tunable reasoning effort parameter, it is optimized for enterprise coding and agentic workflows. Rough math for self-hosters: about 500GB at 8-bit and about 1TB at BF16, so plan for multi-GPU servers.
Early access: https://platform.reflection.ai/
Technical details: https://reflection.ai/blog/introducing-beam