BIG news in AI: Reflection has finally unveiled its much anticipated open weight model, Beam... aiming to be the American answer to China’s open model lead.
As the AI race expands beyond OpenAI and Anthropic, open weight models are surging with developers and moving quickly into the enterprise.
China has dominated with GLM, Kimi and DeepSeek. If Beam can deliver on both capability and efficiency, America may finally have a real contender
https://t.co/6jz21mdeGb
So happy to finally share what we have been working on @reflection_ai . The speed at which the team is building and pushing for the frontier is insane. Cheers to more to come! More on our tech details soon :)
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
- Frontier reasoning efficiency
- Advances the Western open frontier on coding & agentic tasks
- Trained end-to-end from scratch
Full weights release this month.
Learn more about Beam: https://t.co/c3Qx2cpM8G
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
- Frontier reasoning efficiency
- Advances the Western open frontier on coding & agentic tasks
- Trained end-to-end from scratch
Full weights release this month.
Learn more about Beam: https://t.co/c3Qx2cpM8G
Great to see the additive dataset methodology we proposed in Phi-4-reasoning adopted in open-r1.
Tldr: optimize data mixture per reasoning domain, and combine in final run for generalized performance. This is a game changer for reducing data ablation costs.
Happy to share 💭 Mixture of Thoughts 💭
A curated, general reasoning dataset that trims down over 1M samples from public datasets to ~350k through an extensive set of ablations 🧑🍳
Models trained on this mix match or exceed the performance of DeepSeek's distilled models -- not just on math/code but also on scientific benchmarks like GPQA
We also validate that the "additive" methodology from Phi-4-reasoning really works! You can optimise the data mixture independently per reasoning domain and then bring it all together for the final run 🔥
Link to the dataset ⤵️
Excited to share our latest Phi model, Phi4-reasoning, a small but powerful model that match the performance of much larger reasoning models up to DeepSeek R1. Here is the report for new insights into training reasoning models and evaluating them: https://t.co/i7L3utizc0
Introducing Phi-4-reasoning, adding reasoning models to the Phi family of SLMs.
The model is trained with both supervised finetuning (using a carefully curated dataset of reasoning demonstration) and Reinforcement Learning.
📌Competitive results on reasoning benchmarks with much larger top-tier models up to DeepSeek R1
📌 Strong performance on new tests released after data collection (AIME 2025, HMMT)
📌Reasoning transfers/generalizes well to new domains even with only SFT (e.g. k-SAT, Mae Solving, Calendar Planning, etc.)
📌Retains and often significantly improves general-purpose capabilities (e.g. instruction following)
In addition to the models, we are also very excited to share a very detailed technical report with insights on model training and evaluation
Still have a lot to improve especially with context length, coding and tools.
Hope you find the models useful!
A big thanks to the amazing team and to all our partners.
In all, we SFT’ed on ~1.4M reasoning traces on select prompts and further RL'd on a small ~6k sample. Despite the relatively long SFT on select domains, we see broad generalization across domains and no degradation in general purpose performance. On the contrary....🔁📚
Phi-4-reasoning-plus is obtained via a short reinforcement learning on Phi-4-reasoning using a randomly selected subset of SFT prompts. This short RL amplifies the reasoning style and unlocks nice improvements across benchmarks with longer response length.
We’ve been cooking... a new open weights 14B Phi-4 reasoning model, SFT’d on ~1.4M carefully curated reasoning demonstrations from o3-mini and RL’d for a tiny bit. This model is a little beast.
Phi-4-reasoning is supervised fine-tuned on Phi-4. The secret sauce? 1) high-quality prompts at the edge of model capability to go beyond vanilla distillation + strong reasoning responses from a teacher. 2) optimal data mixture of different sources for best overall performance.
More interestingly, our models generalize well to out-of-distribution tasks like algorithmic problem solving, planning, and spatial reasoning. These skills were not targeted in our training data but Phi-4-reasoning performs quite well.
With 14B parameters, both models are competitive and often better than (larger) frontier models: outperforming DeepSeek-R1-Distill-Llama-70B across the board (small gap in coding) and comparable with original DeepSeek-R1 on AIME 2025 which came out after our data cutoff date.
Excited to release our first set of reasoning models Phi-4-reasoning and Phi-4-reasoning-plus, available today on HuggingFace and Azure AI foundry. Some interesting insights below and more deep dives in following days!