Reflection AI Beam is the Nvidia-backed startup’s first model: a 501-billion-parameter, open-weight system designed for coding, reasoning and agentic work. Reflection unveiled the text-only model on October 5, positioning its sparse architecture as a more efficient Western competitor to leading downloadable models from China.

 

The Beam preview has four defining numbers:

  • 501 billion total parameters, with 23 billion active per token.
  • 23.8 trillion tokens used in pretraining.
  • More than 100 million reinforcement-learning rollouts.
  • A one-million-token effective context window.

 

Reflection AI Beam Uses a Sparse 501B Architecture

Beam is a mixture-of-experts model, meaning only a selected part of its network handles each token. Its 23 billion active parameters are less than one-twentieth of the 501 billion stored across the complete model, reducing the computation required for generation while retaining the capacity of a much larger system.

 

The architecture combines fine-grained routed experts with interleaved local and global attention. Reflection says it also uses load-balancing techniques to prevent a small set of experts from receiving most requests, a common optimization problem in sparse models.

 

Reflection trained Beam on 23.8 trillion tokens from web material, public sources and proprietary licensed datasets. The company says its filtering pipeline removed about 95% of raw internet tokens while preserving code, technical explanations and scientific material intended to support downstream agent tasks.

 

Midtraining extended the model’s effective context to one million tokens using code repositories, long-horizon tasks and long-form documents. That specification describes how much material Beam can process, but it is not a guarantee that the model will recall every detail or sustain equal accuracy throughout the window.

 

Beam’s Coding Benchmarks Show Efficiency, Not Overall Leadership

Reflection reports a score of 80.1 on Terminal-Bench 2.1 and 44.4 on DeepSWE 1.1. Beam also reached 80.9 on SWE-bench Verified and 78.7 on MCP Atlas, a tool-use evaluation. Those figures put it near or above several Western open models in the company’s comparison.

 

The same tables show important limits. Beam trails Kimi K3, GLM-5.3 and DeepSeek V4.1 Flash on several agentic coding tests. On Humanity’s Last Exam without tools, its reported 36.2 score is below GLM-5.2, Qwen3.8-Max and Kimi K3.

 

Reflection’s central claim is therefore efficiency rather than absolute benchmark leadership. The company estimates Beam can match GLM-5.2-class reasoning while using three to four times less inference compute, based on active parameters and generated tokens.

 

That estimate excludes prompt processing, attention costs and serving overhead, so it should not be read as a measured price comparison. Independent reporting confirms the launch, but evaluators have not yet reproduced the full benchmark suite using released weights.

 

Background Reading

 

Reflection Scaled Beam With 10,500 Nvidia GB300 GPUs

Beam’s post-training campaign generated more than 100 million rollouts over four weeks on 10,500 Nvidia GB300 GPUs. Reflection used nearly one million coding, terminal, science, web-search and tool-use environments, with a maximum context of 256,000 tokens during reinforcement learning.

 

The company says its asynchronous training system sustained an average of 110,000 concurrent rollouts. Trainers learned from completed agent attempts while inference workers continued generating new ones, and updated weights reached the inference fleet in a median of about 12 seconds.

 

This infrastructure matters because agent training is not just conventional next-token prediction. The model must attempt multi-step tasks, use tools, receive an outcome score and learn from failures. Reflection reports more than 1.3 billion sandboxes were used for training and grading across the main reinforcement-learning run.

 

The scale helps explain why Reflection previously committed billions of dollars to Nvidia-powered capacity from Nebius and SpaceX. Beam is the first public model that links those infrastructure investments to a specific trained system, rather than a promise of future capability.

 

Beam Weights and Safety Results Are Still Pending

Beam is not yet broadly downloadable. Reflection is offering early access to selected users while final red-team testing and evaluations continue. It says the weights, technical report, model card and developer artifacts will be released later in October under the Apache 2.0 license.

 

That timing distinction is material. Developers can evaluate the preview, but they cannot yet audit the promised weights, reproduce the benchmark setup or verify deployment requirements from the complete release package. The technical report will also determine how much detail Reflection provides about data provenance and safety testing.

 

For alignment, Reflection trained a separate safety-focused model and combined it with the capability model through multi-teacher on-policy distillation. Its adversarial curriculum included jailbreaks and scenarios where a tool-using agent was pressured to take unsafe actions.

 

The company has promised to publish safety evaluation results and open-source internally developed evaluations. Until those materials and the weights arrive, Beam’s significance rests on its architecture, training scale and claimed efficiency—not yet on independently verified production performance.

 

If the release lands as described, Beam could give enterprises and governments another route to self-hosted coding agents without relying solely on closed APIs or Chinese open models. The more decisive test will begin after outside developers can reproduce its results, measure real serving costs and examine its behavior on tasks Reflection did not select.