๐ I have joined Reflection AI to lead open source customization stack and the AI factory. After a few great years at Google, I am back to working with a smaller team. Here is why open, why now, and why it is the bet I want to make.
๐ Open models got good this year. Not "good for open," just good. GLM-5.2 is the first open model that feels right in a coding harness as a general agent, and the gap to the closed frontier is now a couple of months, not a generation.
๐ธ The token economics are catching up with everyone. Per-seat pricing is heavily subsidized; move heavy workloads to metered APIs and the invoice tells the truth. Cursor already runs its Composer model on open Kimi weights at ~1/10th of closed-lab pricing.
๐ Then the closed frontier made the case for me. Fable shipped silent degradation, got export-suspended, and forces 30-day data retention. OpenAI previewed GPT-5.6 only to government-vetted partners. When the labs themselves say the access is the problem, believe them.
๐ For an enterprise or a sovereign, a model that can be export-restricted on a Friday, quietly degraded, or made to retain its prompts is a dependency with a kill switch. Nobody wants that sitting under the code they ship.
๐ก๏ธ As the gap closes, the value moves to the data layer. Build on a hosted closed model and you hand your IP, customer data, and metadata to a possible future competitor. Own the stack and you decide where the data goes.
๐งญ Provenance still matters. The Chinese labs do extraordinary research, but a lot of buyers want a model whose training data and values they can trust, open and capable, that they can run themselves. There is a gap here and this is the gap Reflection aims to fill.
๐ฏ My bet: within a year, open models are part of every serious enterprise's production stack, a genuine differentiator. Reflection is building for that, and we are hiring across teams.
Full post here: https://t.co/B7rKOpMvSh
My DMs are open, ping me with any requests for ecosystem that you would like to see alongside the Reflection Models.
@kunchenguid Oh you need to look into the history of this. OpenAI and Google compete to create counter narrative around each others events. I donโt know if this was the case this time. But the early announcement might have been an effort from Google to take some sheen of Dev Day.
@zkwentz from @reflection_ai is the most cracked RL infra person Iโve worked with. Believe me, Iโve worked with a lot.
If youโre in NYC on October 14, go talk to him at Open by Design. Heโll be joining @togethercompute and @nvidia to discuss putting open models into production, post-training for your workloads, and what it costs to get the job done.
Bring hard questions. I promise youโll learn a lot.
October 14 ยท 6โ9 PM ET
https://t.co/8Z4TYN1dpL
I donโt think model pricing tells you much about model quality.
Google/meta can price a model to drive adoption across its products. Another lab may need the API itself to make money. Different hardware and serving stacks also change what they can afford to charge.
A low price could mean a weaker model. It could also mean cheaper inference or a different business objective.
For a customer, the useful comparison is cost per completed task, including the retries and human intervention needed.
Weโre at a point in late stage AI model capitalism where you know the frontier has to win in a bunch of benchmarks to launch. These numbers mean nothing.
The thing to trust is the price.
If they price it high, itโs a good model. If they donโt, itโs benchmaxxed.
Very interesting that the model is SoTA on knowledge work, long context, Harvey LAB etc. Points to the areas of strength for Google and GDM. Looking forward to trying it out.
Introducing Gemini 4 Argon โ our new frontier model.
Itโs built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense โ rolling out today to a set of trusted testers through our Fairwind Program.
@deedydas Depends on whether the taco stand is run by a billionaire (Google/Meta) or someone with a subprime mortgage ๐ . Same taco, very different reasons to charge $0.10 or $10. Price tells you something about the business, too.
@deedydas Rather shallow view no? Especially when more non-NV hardware is being used in inference. Perf/$ across the different chips are quite a bit lower. Including on TPU especially when the company makes them.