New art project.
Train and inference GPT in 243 lines of pure, dependency-free Python. This is the *full* algorithmic content of what is needed. Everything else is just for efficiency. I cannot simplify this any further.
https://t.co/HmiRrQugnP
There are a few thoughts that I have on this:
1. I think we all need to recognize that Anthropic is the only large model provider in the US that actually had such generous, key, unrestricted use. For instance, Codex is effectively unusable outside the original harness. The models are not even all available in the API.
2. Given how Anthropic's API actually works, the underlying unit economics are very dependent on how effective cache utilization is or how the agent, the harness actually drives the loop. So I can totally imagine that when they look at their dashboards, they see wildly differing unit economics from Claude Code where they have a lot of control over it vs what other harnesses are doing.
So I don't think that this is an easy problem and I think it also makes sense for them to have a way to separate out different kinds of uses of the token. But they’re very hard.
I think it's still very hard to find a better deal today. And definitely makes me sad not being able to use this in my favorite way anymore.
I'm not joking and this isn't funny. We have been trying to build distributed agent orchestrators at Google since last year. There are various options, not everyone is aligned... I gave Claude Code a description of the problem, it generated what we built last year in an hour.
yes things are changing fast, but also I see companies (even faang) way behind the frontier for no reason.
you are guaranteed to lose if you fall behind.
the no unforced-errors ai leader playbook:
For your team:
- use coding agents. give all engineers their pick of harnesses, models, background agents: Claude code, Cursor, Devin, with closed/open models. Hearing Meta engineers are forced to use Llama 4. Opus 4.5 is the baseline now.
- give your agents tools to ALL dev tooling: Linear, GitHub, Datadog, Sentry, any Internal tooling. If agents are being held back because of lack of context that’s your fault.
- invest in your codebase specific agent docs. stop saying “doesn’t do X well”. If that’s an issue, try better prompting, https://t.co/SOjpn47yxo, linting, and code rules. Tell it how you want things. Every manual edit you make is an opportunity for https://t.co/S1ZvtYQwta improvement
- invest in robust background agent infra - get a full development stack working on VM/sandboxes. yes it’s hard to set up but it will be worth it, your engineers can run multiple in parallel. Code review will be the bottleneck soon.
- figure out security issues. stop being risk averse and do what is needed to unblock access to tools.
in your product:
- always use the latest generation models in your features (move things off of last gen models asap, unless robust evals indicate otherwise). Requires changes every 1-2 weeks - eg: GitHub copilot mobile still offers code review with gpt 4.1 and Sonnet 3.5 @jaredpalmer. You are leaving money on the table by being on Sonnet 4, or gpt 4o
- Use embedding semantic search instead of fuzzy search. Any general embedding model will do better than Levenshtein / fuzzy heuristics.
- leave no form unfilled. use structured outputs and whatever context you have on the user to do a best-effort pre-fill
- allow unstructured inputs on all product surfaces - must accept freeform text and documents. Forms are dead.
- custom finetuning is dead. Stop wasting time on it. Frontier is moving too fast to invest 8 weeks into finetuning. Costs are dropping too quickly for price to matter. Better prompting will take you very far and this will only become more true as instruction following improves
- build evals to make quick model-upgrade decisions. they don’t need to be perfect but at least need to allow you to compare models relative to each other. most decisions become clear on a Pareto cost vs benchmark perf plot
- encourage all engineers to build with ai: build primitives to call models from all code bases / models: structured output, semantic similarity endpoints, sandbox code execution. etc
What else am I missing?
“Captured decision traces become searchable precedent. And every automated decision adds another trace to the graph.
None of this requires full autonomy on day one.”
This is the DeepSeek moment for Voice AI.
Today we’re releasing Chatterbox Turbo — our state-of-the-art MIT licensed voice model that beats ElevenLabs Turbo and Cartesia Sonic 3!
We’re finally removing the trade-offs that have held voice AI back.
Fast models sound robotic. Great models are slow. And none are built for trust. We fixed all three. Chatterbox Turbo is transparent, auditable, and built for a world that needs proof.
There are no limits anymore. Anyone can do anything. The only limiting factors are agency and ambition.
Never has a college degree, work experience, network, even the accumulation of knowledge been worth less.
You can just ship things.
.@flowglad is an open-source, zero-webhook payments provider.
Their full-stack SDK lets you check customers' features and usage credits in real time, based on billing state.
And their AI generates a setup prompt for your pricing model and stack, so you can one-shot payments.
https://t.co/8ETTLBFW0u
A number of people are talking about implications of AI to schools. I spoke about some of my thoughts to a school board earlier, some highlights:
1. You will never be able to detect the use of AI in homework. Full stop. All "detectors" of AI imo don't really work, can be defeated in various ways, and are in principle doomed to fail. You have to assume that any work done outside classroom has used AI.
2. Therefore, the majority of grading has to shift to in-class work (instead of at-home assignments), in settings where teachers can physically monitor students. The students remain motivated to learn how to solve problems without AI because they know they will be evaluated without it in class later.
3. We want students to be able to use AI, it is here to stay and it is extremely powerful, but we also don't want students to be naked in the world without it. Using the calculator as an example of a historically disruptive technology, school teaches you how to do all the basic math & arithmetic so that you can in principle do it by hand, even if calculators are pervasive and greatly speed up work in practical settings. In addition, you understand what it's doing for you, so should it give you a wrong answer (e.g. you mistyped "prompt"), you should be able to notice it, gut check it, verify it in some other way, etc. The verification ability is especially important in the case of AI, which is presently a lot more fallible in a great variety of ways compared to calculators.
4. A lot of the evaluation settings remain at teacher's discretion and involve a creative design space of no tools, cheatsheets, open book, provided AI responses, direct internet/AI access, etc.
TLDR the goal is that the students are proficient in the use of AI, but can also exist without it, and imo the only way to get there is to flip classes around and move the majority of testing to in class settings.
We just published a new AI-Native Engineering Team guide based on what engineering teams are asking for as they adopt Codex and the new GPT-5.1-Codex-Max model.
It covers:
🧩 How coding agents fit into each phase of dev across planning, design, maintain
🧰 Practical checklists and setup patterns you can use right away
📈 How to introduce agents into an org and scale as teams build trust
Read the guide 👉 https://t.co/yyM3aIF2Vo
GitHub repo:
https://t.co/Cpm3Dc44rY
A lot more detailed and technical walkthrough:
https://t.co/YmHaZfNjcJ
Example conversation with the $100, 4-hour nanochat in the WebUI. It's... entertaining :) Larger models (e.g. a 12-hour depth 26 or a 24-hour depth 30) quickly get more coherent.
There’s going to be a huge gap in execution velocity for the foreseeable future between the individuals, teams, and companies that adapt (or create for the first time) their workflows to work with AI agents vs. those that don’t.
In the past few months I’ve talked with a number of new startup founders who are operating in a totally different way than the rest of the world right now.
Most of their work, particularly in engineering to start with, is oriented around how to make agents effective. There’s a focus on hyper specific prompting, a bigger orientation around getting specs perfectly right, running many agents in the background in parallel, focusing on code reviews vs. coding, and a bunch of other new workflow practices are actually what it takes to make agents work at scale.
And while coding is ahead of the curve right now in agentic workflows, it’s clear this pattern will start to emerge across most other domains over time. Leverage is going to show up everywhere: the ability to generate 10X more marketing output, process deals way faster due to automated legal workflows, handle customer support and success operations in faster ways, and so on.
The lesson here is that these teams just tend to be far more ambitious in what they push agents to do. Most existing teams and companies will be happy about incremental gains and stop short of doing the really big changes in how they work. And doing so is going to provide at least a temporary advantage to those that adapt to these news ways of working.
In era of pretraining, what mattered was internet text. You'd primarily want a large, diverse, high quality collection of internet documents to learn from.
In era of supervised finetuning, it was conversations. Contract workers are hired to create answers for questions, a bit like what you'd see on Stack Overflow / Quora, or etc., but geared towards LLM use cases.
Neither of the two above are going away (imo), but in this era of reinforcement learning, it is now environments. Unlike the above, they give the LLM an opportunity to actually interact - take actions, see outcomes, etc. This means you can hope to do a lot better than statistical expert imitation. And they can be used both for model training and evaluation. But just like before, the core problem now is needing a large, diverse, high quality set of environments, as exercises for the LLM to practice against.
In some ways, I'm reminded of OpenAI's very first project (gym), which was exactly a framework hoping to build a large collection of environments in the same schema, but this was way before LLMs. So the environments were simple academic control tasks of the time, like cartpole, ATARI, etc. The @PrimeIntellect environments hub (and the `verifiers` repo on GitHub) builds the modernized version specifically targeting LLMs, and it's a great effort/idea. I pitched that someone build something like it earlier this year:
https://t.co/ANHhasxzD8
Environments have the property that once the skeleton of the framework is in place, in principle the community / industry can parallelize across many different domains, which is exciting.
Final thought - personally and long-term, I am bullish on environments and agentic interactions but I am bearish on reinforcement learning specifically. I think that reward functions are super sus, and I think humans don't use RL to learn (maybe they do for some motor tasks etc, but not intellectual problem solving tasks). Humans use different learning paradigms that are significantly more powerful and sample efficient and that haven't been properly invented and scaled yet, though early sketches and ideas exist (as just one example, the idea of "system prompt learning", moving the update to tokens/contexts not weights and optionally distilling to weights as a separate process a bit like sleep does).
I am (slowly) re-reading the Tolkien legendarium (of which Lord of the Rings is a small part). The whole body of work is so incredible and there's nothing else like it... it dilutes other worlds of fiction. Wait - your story doesn't have a comprehensive history/mythology spanning multiple ages all the way back to a creation myth as detailed in separate volumes? You didn't first invent new languages and dialects for your characters? You didn't pack it with powerful themes and stories written it in a beautiful, archaic style and compose poems and songs alongside? It didn't take you multiple decades of iteration? And what of all the uncharted territory still remaining? Is Tom Bombadil one of the Ainur. Where are the Entwives. What happened to the two unaccounted Istari. Can we hear more about what it was like in Cuiviénen when the elves first awoke? Or to see the light of the two trees of Valinor. Or of the splendor of the caves of Aglarond.
What's most on my mind though - the Tolkien legendarium is imo a concrete example of a height of culture. Does AI, today or soon, make it easier to reach this high via empowerment in both writing and ideation? Or harder, when quick wins are tempting and ~free, and an independent ability to create is stifled. If such a body of work is made again but now with heavy AI assistance, does it inspire the same wonder? What if thousands of them come out on demand with just a prompt? Why do you feel cheated when you learn that something your read was AI generated? Is it transient or a function of capability? Is it slop? What is slop? Or is wonder inseparable from its own creation myth of a lifelong obsession of a mind like your own? So many questions.
Working on Torch 🔥 Unified health record + LLM in one iOS app.
Syncs records from hospitals, labs, Function, One Med, PDFs etc. Makes it simple & fast to get all your health data in one place + get LLM help making sense of it.
Reply/RT for TestFlight invite. More below ⬇️