Reflection AI, the Nvidia-backed American startup, introduced its first open-weight model on October 5, a sparse mixture-of-experts system called Beam that the company positions as the strongest Western challenge yet to the open-weight frontier currently dominated by Chinese labs. The announcement, published on Reflection's own blog and quickly picked up by Reuters, TechCrunch, Fortune and Semafor, arrives with a twist: the model is not downloadable yet.
Beam is undergoing what the company describes as final red-teaming and evaluations. Early access is open through a signup process, and Reflection says it will release the weights, a technical report, a model card and developer artifacts later this month. When that happens, Beam will become one of the largest open-weight releases ever published by a US lab, and the clearest test yet of whether an American company can match the release cadence and openness that has made Chinese models the default choice for much of the open-source community. For continuous coverage of the open-weight race and everything else moving the field, bookmark our AI news homepage.
The Numbers Behind Beam
Beam is a sparse mixture-of-experts model with 501 billion total parameters, of which roughly 23 billion are active per token. Reflection says it pretrained the model on 23.8 trillion tokens drawn from the web and proprietary licensed datasets, claiming that the base model matches or outperforms similar-sized open base models.
The more distinctive claim concerns reinforcement learning. Reflection built the algorithms, training environments and infrastructure to sustain high-compute RL at scale, and the company reports that its RL run generated more than 100 million rollouts across 10,500 NVIDIA GB300 GPUs over four weeks of training. That is an unusually large RL investment for an open-weight release, and Reflection credits it for Beam's competitive performance with what it calls frontier inference compute efficiency.
Benchmarks: Competitive, Not Yet Leading
Reflection's published benchmark table places Beam solidly in the open-weight pack, behind the very largest Chinese models but ahead of or level with several established names.
On SWE-Bench Pro v2-Hard, Beam scores 77.2, behind GLM 5.3 at 88.2 and Kimi K3 at 88.2 on the same row, but well ahead of Inkling at 56.9. On SWE-Bench Pro v1, Beam scores 65.5, ahead of Inkling (54.3), Nemotron 3 Ultra (46.4) and GLM 5.2 (62.1), while Qwen 3.8-Max posts 67.7. On the DeepSWE v1.1 agentic coding benchmark, Beam records 44.4, level with GLM 5.2's 44.0 and behind GLM 5.3 (61.0), Qwen 3.8-Max (51.0) and DeepSeek V4.1 Flash (74.2).
The company's own framing is candid about where Beam sits. In the blog post, Reflection writes that Beam "advances the Western open-weight frontier" and is "competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks." It also concedes that frontier open models like Kimi K3 "remain ahead on raw capability."
Two caveats deserve emphasis. First, every number here is self-reported by Reflection ahead of the weight release, and independent verification will only become possible once the technical report and checkpoints are public. Second, several cells in the company's own table are marked as not reported, which makes complete apples-to-apples comparisons impossible for now.
The Efficiency Argument
What Reflection pitches as Beam's real advantage is not raw leaderboard position but cost at inference time. Because only 23 billion of its 501 billion parameters activate per token, the company argues Beam can deliver near-frontier agentic and coding performance at a fraction of the compute cost of larger dense models. That positioning mirrors the strategy that made Chinese open-weight families so popular with startups: strong-enough capability at prices that make always-on agents economically viable.
If the efficiency claim survives independent benchmarking, it could matter more than leaderboard deltas. Teams running thousands of parallel coding agents care more about cost per completed task than about single-digit benchmark points.
A Geopolitical Undercurrent
Coverage of the launch has been strikingly geopolitical for an open-source model release. Reuters described Beam as the "first AI model" from Reflection to "take on Chinese open models." Fortune asked whether Beam could be "America's best chance to compete with China," and Semafor framed the company as "an open-source Western answer to Chinese labs."
That framing reflects the state of the open-weight ecosystem in late 2026: the most capable downloadable models have come overwhelmingly from Chinese labs including Z.ai's GLM family, Moonshot's Kimi line, Alibaba's Qwen and DeepSeek. American labs have largely kept their best weights closed, betting on API revenue instead. Reflection is betting the opposite: that an open Western frontier model can attract the developer ecosystem, and that Nvidia's backing gives it the compute runway to keep up.
What Happens Next
The weight release later this month will answer the questions that matter: how Beam performs outside Reflection's own evaluations, what license accompanies the weights, and whether the efficiency holds up under independent load. Until then, Beam remains a promising benchmark table, a signup form, and the most consequential open-weight promise an American lab has made this year.
For developers deciding whether to standardize on GLM, Kimi, Qwen or wait for Beam, the practical advice is to hold final decisions until the technical report and third-party evaluations land. The open-weight race has a new American entrant, and for the first time in a while, it is not starting from the back of the pack.
Stay Ahead of AI
The open-weight frontier moves weekly, and the labs release faster than anyone can track. Read more AI news on AI Buzz Wire for model launches, benchmark breakdowns and the business stories behind them.
Explore the latest AI developments here.