The model
Reflection is a US company founded in 2024 by two former Google DeepMind researchers, which has raised about $4.7 billion. On October 5 it introduced Beam, its first open-weight model: a sparse mixture-of-experts design with 501 billion total parameters and 23 billion active, pretrained on 23.8 trillion tokens with a 256K training context extended to an effective 1 million tokens. Reported training compute: 6,144 Nvidia GB300 GPUs for under four weeks of pretraining, and about 10,500 GB300s for four weeks of reinforcement learning, generating over 100 million rollouts.
You can't download it yet. Weights are due later this month under Apache 2.0, along with a technical report, model card and fine-tuning tooling; the model is still in final red-teaming and evaluation, with only an early-access sign-up for now. Training data and the full training pipeline are not part of the release.
Read the whole scorecard
Reflection reports 80.9% on SWE-Bench Verified and 97.8% on AIME 2026. But on the same table it doesn't lead: DeepSWE v1.1 is 44.4%, against 68.0% for Kimi K3 and 51.0% for Qwen 3.8 Max; Terminal Bench v2.1 is 80.1%, below Kimi K3's 88.3%; and GPQA Diamond, at 90.5%, is the lowest in the table. Reflection itself acknowledges that frontier open models like Kimi K3 remain ahead on raw capability, and pitches Beam on inference efficiency: comparable to GLM 5.2 on advanced reasoning benchmarks with 3–4x less inference compute. That estimate excludes prefill, context-dependent attention operations and serving overhead, so it is an approximate comparison, not a measured cost.
What it means
The top tier of open-weight models is currently held by models from Chinese teams, and Beam is one of the few Apache 2.0 options of that size from a US company, which matters to enterprises with data compliance or vendor-origin requirements. Until the weights actually ship and third parties reproduce the results, though, every number is self-reported. Teams eyeing it for coding agents should watch the DeepSWE gap closely and test on their own tasks when it lands.
via: Reflection blog, TechCrunch report, MarkTechPost report