🚀 I have joined Reflection AI to lead open source customization stack and the AI factory. After a few great years at Google, I am back to working with a smaller team. Here is why open, why now, and why it is the bet I want to make.
📈 Open models got good this year. Not "good for open," just good. GLM-5.2 is the first open model that feels right in a coding harness as a general agent, and the gap to the closed frontier is now a couple of months, not a generation.
💸 The token economics are catching up with everyone. Per-seat pricing is heavily subsidized; move heavy workloads to metered APIs and the invoice tells the truth. Cursor already runs its Composer model on open Kimi weights at ~1/10th of closed-lab pricing.
🔒 Then the closed frontier made the case for me. Fable shipped silent degradation, got export-suspended, and forces 30-day data retention. OpenAI previewed GPT-5.6 only to government-vetted partners. When the labs themselves say the access is the problem, believe them.
🔌 For an enterprise or a sovereign, a model that can be export-restricted on a Friday, quietly degraded, or made to retain its prompts is a dependency with a kill switch. Nobody wants that sitting under the code they ship.
🛡️ As the gap closes, the value moves to the data layer. Build on a hosted closed model and you hand your IP, customer data, and metadata to a possible future competitor. Own the stack and you decide where the data goes.
🧭 Provenance still matters. The Chinese labs do extraordinary research, but a lot of buyers want a model whose training data and values they can trust, open and capable, that they can run themselves. There is a gap here and this is the gap Reflection aims to fill.
🎯 My bet: within a year, open models are part of every serious enterprise's production stack, a genuine differentiator. Reflection is building for that, and we are hiring across teams.
Full post here: https://t.co/B7rKOpMvSh
My DMs are open, ping me with any requests for ecosystem that you would like to see alongside the Reflection Models.
@Voxyz_ai I am a big believer in this pattern. I don’t think folks have truly embraced multi model workflows yet. Most people I know run the best model (Astra, Fable) in ultra mode and then complain about their quota running out. 😂
It seems like the entire industry conspired to make sandboxes and cloud managed agents as the thing to be working on and launching this week.
How long before we have sandbox “benchmarks” or do they already exist?
Today at @WeAreDevs, Docker and the @linuxfoundation are announcing a collaboration around the Docker Sandbox Kit Specification: an open standard for declaring what an agent may do, where it may reach, and what it may touch.
Open source under Apache 2.0.
Read the deep dive: https://t.co/meuAmgy5Ln
I would think this categorization os highly related to the task right? Do you have a summary of the type of tasks your usually run?
Also, I see that you are not that into open models? Is the intelligence gap compared to the price not that attractive or is the reason something else?
@achowdhery In steady state having a cascade of models and smart routing would make this a much more palatable. Most tasks that people are doing doesn't need a model bigger than 100B. But I am not sure when we will get to steady state. 🤷♂️
Releasing SmolDataEnvs 🤗
5K+ Verifiable RL Environment tasks for hill-climbing small models in code and data science.
Completely open source: environments, evals, training
@trq212 As a PM who likes to prototype, I find it very useful to solidify my ideas and make it more concrete. I am not always familiar with the codebase and very often have ideas that are not fully fleshed out. Haven’t tried artifacts for this, will give it a try.
@ArtificialAnlys Love the animation. Cost per task is the right metric. However, we need a real world task completion index. Even with the best of intents all models eventually are benchmaxxed.
Go talk to @joespeez (VP of Product at @reflection_ai) at Sierra Ventures’ CXO Summit this week. He’s going to be on a panel with @zedlewski talking open-weight models and building ecosystems. Ask him how you can adopt Open Models for your own team/organization.
See the quoted text from the podcast: "And we can get to $5 billion in revenue by going deep across those things and creating the systems of record that help manage them. That for Anthropic would be like stopping on the side of the road to pick up a penny, because they're on the pathway of trying to go from $100 billion in revenue to a trillion in revenue."
@levie has also made this argument previously. Getting enterprises in each vertical to adopt AI requires deep understanding of the space and embedding FDEs and other resources to get it done right.
@GabeStengel makes a great point about building moats in specific verticals that can compete with the frontier labs’ offerings. Open-weight models give Rogo one more layer of differentiation: the ability to customize the model itself for how these firms actually work. Better quality and lower inference costs will improve customer outcomes while improving Rogo’s margins.
Gabe on why he doesn't think Anthropic or OpenAI will own finance:
"The reason people get confused when they look at app layer businesses like mine or Harvey or Legora or Sierra is because there's a spectrum of perpendicularity to what the labs are building.
There's a whole bunch of stuff underneath the surface that the labs are never going to build, that we need to build for finance.
All of finance is a collection of different niches with different data sets, different regulatory requirements.
And we can get to $5 billion in revenue by going deep across those things and creating the systems of record that help manage them.
That for Anthropic would be like stopping on the side of the road to pick up a penny, because they're on the pathway of trying to go from $100 billion in revenue to a trillion in revenue.
Say you are a big public company buying another big public company and you need to send data back and forth.
You actually need some sort of data room, something that is compliant, safe and secure. And I don't think OpenAI or Anthropic will ever want to build a data room business.
If you actually want to be the exchange for all of high finance, you don't just need to own the intelligence, you need to own the transaction venue, the communication venue, the workflows, and all the data inputs that go into it.
Think about the fundamental difference between Claude Code when it came out versus ChatGPT.
The models were actually fairly similar, but the harness and the way that it was presented from Claude was far better.
The way that you harness these models is so, so important."
@levie Jevons paradox means cheaper AI will lead to more AI use. So how do we tell whether we're making meaningful progress, rather than just consuming more?
Is cost per successfully completed task the right measure? If so, have you figured out a way to measure it.