Reinforcement Learner for Life
Original proponent of the concept of Agent Trajectory & Recursive Self Improv
I knowledgeGraph anything
Firm & Fluid Believer
No one (actually) working on open source is surprised by this release. For months, @reflection_ai engineers have been down in the trenches contributing to projects like transformers, vLLM, SGLang, TRL, and OpenEnv. Go check the commits.
I'm super excited to see an open lab form an identity for itself that is low on tweets, high on commits; low on tokenmaxxing, high on efficiency.
As an MLE working in open source, it is so fucking cool that companies like this exist, and we all just need to soak in what is coming.
@lovesosaa888@HarveenChadha agreed...openai and anthropic looks like have more chinese folks than proabbly even the chinese labs themselves 😂 , anytime u see their demo videos u will see them more often than not..
We are introducing Beam today, a 500b open model that excels in coding, agentic, and scientific workloads.
Beam is very token-efficient, pre-trained from scratch, and scaled with the largest RL run documented openly we are aware of.
Getting here was challenging, rewarding, and personal.
I immigrated to the USA during childhood when my parents got jobs as scientists at PNNL, a national lab located in a rural town in Washington state.
Growing up, I developed an interest in physics because it was grounded in simple principles, was foundational, and had long-term impact.
Physics was the foundational science of the 20th century. The first general computer (ENIAC), the transistor, GPS, and much of modern technology had its roots in physics.
In 2016, during my graduate studies, I noticed the birth of a new foundational science - Artificial Intelligence.
AlphaGo, the superintelligent Go agent developed by DeepMind, had come out and made me feel what was about to come.
A neural network, trained to imitate human players, and then improve itself through trial and error, mastered the most intellectually challenging board game in the world.
What surprised me about AlphaGo is how it was conceptually simple, general, and scalable. The system could in principle improve itself forever, it was just a function of compute.
This is when I became “AGI pilled” as they say.
I switched from physics to AI because I thought this would become the root node science of the 21st century, much like physics was the root node science of the 20th century.
I believe that working on foundational science is not enough on its own. Foundational science must also remain open and widely accessible.
We remember Newton for his contributions to physics, but he was also an alchemist.
Alchemy was done in secret, because the artificial production of gold was deemed lucrative but also dangerous with the potential do destroy entire economies. Newton even wrote his alchemy notebooks in cypher.
On the other hand, Principia was published in the Philosophical Transactions of the Royal Society, the knowledge was made widely available, passed down for generations in what we now call Newton’s laws.
The computer, transistor, GPS were possible to build because physics as a science was open. Openness creates an ecosystem of innovation.
The notion of openness, which is a core tenet of scientific discovery, has also become a guiding principle for how foundational technology gets built and distributed.
The internet was built on open protocols. The most widely used operating systems are open. The field of cybersecurity exists because strong encryption protocols were made open in the 1990s.
Open also means safer. The best defence is a distributed one. Commonly used open source software is robust and battle-tested because it is inspected and patched by a community of developers.
AI is different but also similar to software in many ways, though with a much larger surface area of vulnerabilities. It is impossible for a small group of safety researchers working in secrecy to remediate all of the potential issues regardless of their intentions. It is like entrusting a handful of white blood cells to protect your body.
AI is both foundational as a science and a technology, and the safest way that also ensures the benefits of AI are distributed evenly is by making it open.
I’m excited for what developers and scientists build and learn from this model, and for Beam to play a small part in ensuring that AI is made widely accessible, safe, and beneficial to all.
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
- Frontier reasoning efficiency
- Advances the Western open frontier on coding & agentic tasks
- Trained end-to-end from scratch
Full weights release this month.
Learn more about Beam: https://t.co/c3Qx2cpM8G
We built our first model!! Beam is a competitive text-only 501B-A23B MoE, openly released under the Apache 2.0 license
It’s been fun scaling the production RL/OPD stack for this over the past year. I’m so proud of the team! We attained exactly what we set out to do and ended up with a stable run on 10K GB300s that produced more than 100M rollouts across ~1M tasks that significantly improved the final performance of the model
👇 many more details in our announcement below :)
Artificial Analysis has been given access by Reflection and is independently benchmarking Beam
Early indicators suggest Beam will be one of the most token-efficient open models we've seen for its level of intelligence.
Congratulations @reflection_ai on the announcement!
I am so proud of Beam.
Beam is Reflection's first open-weight model. It's the result of a ~year of hard work, ingenuity, and camaraderie of one of the most incredible teams I've ever seen in AI 🔥
I joined Reflection last November. Pretraining in particular was so nascent – only 5-10 people and a bit of code! We didn't yet have a crawler, or a resilient GPU infra, or a deduplication pipeline, or in-house evals, or even sufficiently big MoEs. All that started in earnest in 2026, in one of the most satisfying sprints I've ever had the pleasure to live through.
What we nailed, I think, was a perfect blend of scientific rigor, startup intensity, team iteration, and ambition. We did not YOLO decisions in pretraining, that tends to blow up in your face when you scale up 🙂 But we also made sure to move fast and play to our strengths as a nimble, high-ownership, high-trust team. The resulting system of systematic iterations and fast cross-team back-and-forths really paid off.
The midtraining, RL, and post-training teams really really cooked here on top of the pretrained base. The announcement below details the sheer scale and payoff of that investment, I'll let the team cover it, it's remarkable. On many agentic coding capabilities the model kept going up without any sign of a plateau.
Feel free to read more in the announcement today – but later this month we will also drop the full weights release under Apache 2.0, a detailed tech report with the cool LLM science, and partner integrations with the OSS ecosystem. It's an open model after all. I strongly believe open intelligence is the future of AI, this is why I'm here.
And I encourage y'all to join us on the open side 😉
Beam combines strong agentic performance, efficient reasoning, and a 500B form factor to give enterprises, governments, and developers a true workhorse open model. It's our first step toward building frontier intelligence that is open and widely accessible to all.
Now in the final stages of red-teaming, Beam will be released under an Apache 2.0 license this month, along with quantized FP8 and NVFP4 numerics for efficient deployment.
Sign up for early access: https://t.co/z7GJX2ejdN
so many great open research avenues in multi-modal around visual perception/reasoning, context management, tool use from classical CV models, etc
also always have a soft spot for vision from phd & startup i did haha
https://t.co/IkNqKxkMJY
here's the link - not directly related to agent stuff, a more up to date version of this paper would add dinov3 pre-built in registers + some visual agentic task after. but retrieval research alone is still interesting!