Reflection AI has announced Beam, an open-weight model with 501 billion total parameters and 23 billion active parameters. In a post on X, @reflection_ai describes it as an agentic model aimed at coding, reasoning and other multi-step tasks, while saying its full weights are planned for release under the Apache 2.0 license during the month of the announcement.
The company also reports performance and inference-efficiency comparisons for Beam. Those results are part of Reflection AI’s launch announcement, and the posts do not include the full evaluation methods, serving conditions or release documentation needed to independently assess them.
What Reflection AI announced about Beam
Beam’s headline architecture figures are 501 billion total parameters and 23 billion active parameters. Parameters are the learned values that store patterns in a model. The total count describes the model’s full parameter capacity, while the active count describes the parameters used for a particular computation when the model uses a sparse design.
That distinction matters because a model can have a large overall capacity without applying every parameter to every input. A smaller active count can reduce the computation required for individual requests compared with running the full parameter set each time. Actual speed and cost still depend on hardware, serving software, workload and other deployment choices.
Reflection AI says Beam was trained end-to-end from scratch and positions it as an open model for developers, enterprises and governments. The company describes its focus as coding and agentic tasks. In this context, “agentic” refers broadly to systems that can carry out multi-step tasks, but the announcement does not define a specific agent evaluation protocol.

Image credit: @reflection_ai on X
Beam’s reported efficiency and capability position
Reflection AI says Beam leads its model class in inference efficiency. The company reports that it is three to four times more efficient than GLM 5.2 and more than four times more efficient than leading Western open models, adding that Beam completes tasks faster and more cheaply.
Those comparisons remain company-reported results. Reflection AI’s posts do not specify the hardware, batch size, workload, latency or cost definitions, serving stack or evaluation methodology behind the efficiency figures. The reported multiples therefore describe the company’s position for Beam rather than an independently established ranking.
One of Reflection AI’s graphics compares Beam with several other models across DeepSWE v1.1, Terminal Bench v2.1, HLE No Tools, SWE Bench Pro V1, SWE Bench Verified and CritPT AA. The graphic marks some entries as “NR,” meaning “Not Reported.”
The displayed Beam scores are 44.4 on DeepSWE v1.1, 80.1 on Terminal Bench v2.1, 36.2 on HLE No Tools, 65.5 on SWE Bench Pro V1, 80.9 on SWE Bench Verified and 15.6 on CritPT AA. The comparison does not show that every model was tested under identical conditions, and the announcement does not provide the prompts, sampling settings, dates or confidence information needed to interpret the figures as a complete independent evaluation.

Image credit: @reflection_ai on X
A second graphic plots Terminal Bench scores against estimated generation-forward FLOPs per attempt. It presents Beam as part of an efficiency comparison, but the announcement does not define how those estimated compute figures were calculated or whether they represent serving cost, latency or another measure.
How Reflection AI says Beam was trained
Reflection AI says Beam was pretrained for four weeks on 24 trillion high-quality tokens. The company says the pretraining run was designed to give the model innate coding capabilities and credits data curation, deduplication and innovations in mixture-of-experts stability for its foundation.
A mixture-of-experts, or MoE, model uses specialist parameter groups rather than applying the entire parameter set to every input. That approach can help explain why a model may have a much larger total parameter count than active count, although the announcement does not describe Beam’s routing design, number of experts or other implementation details.
The company also reports a reinforcement-learning run using 10,500 GB300 GPUs for four weeks. Reinforcement learning is a training stage that uses feedback to improve a model’s behavior on selected tasks. Reflection AI says its algorithmic and distributed-infrastructure work allowed this process to scale, and that capabilities continued to improve as reinforcement learning increased without signs of a plateau.
The company characterizes this as the largest publicly documented reinforcement-learning run it is aware of. That description is an attributed company assessment, not an independently established comparison. The announcement does not publish the underlying evaluation data or a reproducible account of the scaling experiment.

Image credit: @reflection_ai on X
The announcement also includes graphics comparing Beam Base with other base models on code and web validation loss against training compute. The charts are presented as evidence for Reflection AI’s description of Beam’s pretraining, with lower validation loss shown as better. They do not include enough methodological detail to establish how the comparisons were produced or how they should translate to deployed-model performance.

Image credit: @reflection_ai on X
What the planned open release is expected to include
Reflection AI says Beam was in the final stages of red-teaming when it announced the model. Red-teaming is a process for probing a system for safety problems, failure modes and misuse risks. The company said the weights would be released during the month of the announcement under the Apache 2.0 license, along with FP8 and NVFP4 numerical formats for more efficient deployment.
FP8 and NVFP4 refer to lower-precision numerical formats that can reduce memory and computation requirements when the supporting hardware and software stack can use them. The announcement does not include the model files, final license text, a model card, context-window information, hardware requirements, safety documentation or installation instructions.
The release timing is also prospective. Because the announcement does not identify its month or provide a specific calendar date, the promised release window cannot be converted into a confirmed date. The source does not establish that Beam’s weights are already downloadable or ready for production use.
What Beam could mean for developers and enterprises
If the reported architecture and efficiency results hold under reproducible testing, Beam’s large total parameter count paired with a smaller active count could interest users who want broad model capacity without applying all parameters to every request. The practical value would depend on the eventual weights, supported hardware, inference software, memory requirements and measured performance on users’ own workloads.
A final Apache 2.0 release could make the model easier to inspect, adapt or deploy than a hosted-only system, subject to the final license terms and release materials. Questions about commercial use, safety restrictions, fine-tuning and operational requirements remain open until those materials are published.
For now, Beam is best understood as Reflection AI’s announced open-weight model and its accompanying performance and efficiency proposal. Independent evaluations and release documentation will be needed to determine how the 501-billion-parameter design performs in practice.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment