Reflection AI just introduced Beam, its first open-weight model. The headline number is 501 billion parameters. The number builders should actually care about is 23 billion.
What was announced (Oct 5, 2026)
- Architecture: a sparse Mixture-of-Experts (MoE) model with 501B total parameters and about 23B active per token, built for coding, reasoning, and agentic work.
- Pretraining: 23.8 trillion curated tokens, trained in under four weeks on a GB300 NVL72 cluster.
- Reinforcement learning at scale: 10.5K NVIDIA GB300 GPUs for four weeks, more than 100 million rollouts, about 1.3 billion sandboxes, and close to one million coding, agentic, and STEM environments.
- Context: midtraining extends effective context to 1M tokens.
- License and timing: early access now. Reflection says weights, a technical report, a model card, and fine-tuning tooling ship later this month under Apache 2.0.
Expected vs. actual
Expected: a bigger open model means more hardware and a bigger bill for every answer.
Actual: in an MoE model, only a slice of the network wakes up for each token. Reflection reports Beam scores comparable to GLM-5.2 on advanced reasoning benchmarks while using roughly 3 to 4 times less inference compute, and says it approaches Qwen 3.8-Max on coding and agentic tasks. (These are the company's own benchmarks, so wait for independent evals before you bet a roadmap on them.)
A plain way to picture it
Think of a hospital with 500 specialists on staff. You don't see all 500 for a fever. A triage desk sends you to the two or three doctors who fit. Total staff tells you how much the hospital knows. The doctors in the room tell you what the visit costs. MoE works the same way: total parameters are knowledge, active parameters are the bill.
Three things worth noticing as a builder
- "Intelligence per token" is the new spec sheet. Beam has a reasoning-effort setting, trained with a length penalty, so you can trade answer depth for tokens per request. That's a direct cost lever for agent loops.
- RL is now a scaling axis, not a finishing step. Reflection says capabilities kept improving with more RL compute with "no sign of a plateau", and browsing improved even though browsing wasn't in the RL mix.
- Open weights plus Apache 2.0 changes the build-vs-buy math. If the release lands as promised, teams get a permissive, self-hostable coding and agent model from a US lab, at a time when most top open models come from Chinese labs.
What I'd do this month
- Join the early-access list if you run coding agents or MCP-heavy workflows.
- When the weights drop, benchmark Beam on your repo tasks, not just public leaderboards.
- Track cost per solved task, not cost per token. A model that solves the task in fewer tokens wins even at a similar per-token price.
References
- Reflection AI, "Introducing Beam: Reflection's 501B open-weight model" (Oct 5, 2026): https://reflection.ai/blog/introducing-beam
- SiliconANGLE, "Reflection AI debuts open-source Beam model with 501B parameters" (Oct 5, 2026): https://siliconangle.com/2026/10/05/reflection-ai-debuts-open-source-beam-model-with-501b-parameters/
- Wang et al. (DeepSeek-AI), "Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts" (cited by Reflection): https://arxiv.org/abs/2408.15664
Benchmark claims above are Reflection AI's own reported numbers.
Top comments (0)