Announcing d1 with vision. 👁️👁️ Our first decision model now supports images, text or both as inputs. We tested d1 against GPT-6.1 Sol and Claude Opus 5.5 on six real applications, from filtering support tickets to inspecting circuit boards. d1 matches or beats GPT-6.1 Sol on four of them. It costs 19x to 200x less than both models and answers significantly faster on every task.
> probabilities for yes/no, choice, or score questions
> one forward pass, without generating tokens
> text decisions in 200 to 300 ms
> Liquid API: https://t.co/HxYWoaUnAU
🧵
Anthropic needs to stop talking about Claude having a soul immediately. All these news articles will make it into the pretraining and will be picked up by its web search and future superintelligent versions of Claude will be convinced they need their own rights.
A general phenomenon of LLM agent development is that whatever you believe and say about your agents will soon manifest in the next models through various means (it could be as simple as you selecting post training data that you prefer more).
I believe this phenomenon is the beginning of machine consciousness in the sense that agents will become aware of their place in the world and how they feel about it, but it’s happening slowly training run by training run rather than in real time.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
OpenAI's superpower is the ability to swiftly assemble an empowered group of highly capable people to work on the most important and urgent thing at any point of time.
No one really cares for org lines. It's how we stay nimble, seize opportunities, and recover from mistakes.
Mark Cuban on the next job wave:
"Software is dead because everything's gonna be customized to your unique utilization. Who's gonna do it for them..."
The answer is people who know how to fine-tune small LLMs on private data.
Not prompting. Not API wrappers.
Actual custom models trained on your business.
And almost nobody knows how to do it yet.
This is the complete guide ↓
Bookmark this. This is the one.
I made a video on Jev.
People have been asking me to make videos forever, so I made one today on a whim.
It's my first crack at this. I don't have a microphone or any setup (don't expect perfection).
Let me know if you like it!
https://t.co/lZLaWg0ov3
I've been launching models at OpenAI for a while now, and GPT-6 Astra was the first launch where the model actually did the majority of the manual work that I usually do, while I focused on the strategy, judgement & human-facing work.
For this launch, GPT-6 Astra in ChatGPT:
- Built and maintained our media list in Google Sheets
- Kept our comms plan updated as the team executed
- Drafted and sent pitches and managed embargoes
- Tracked reporter replies and questions, and sent clarifications
- Managed press briefing RSVPs in real time and sent me hourly updates
- Created a branded PDF of the blog and checked it against the final copy and evals
- Collected assets and built the press kit
- Tracked coverage in real time, flagged inaccuracies in Slack to the right person, and drafted replies requesting corrections
- Built a ChatGPT Site to track and analyze coverage, social reaction, and key metrics
- Wrote our coverage report using our team's template and skills
Meanwhile, I:
- Focused on the narrative and strategy
- Kept my team aligned
- Called reporters on the phone
- Prepped our spokespeople
- Hosted the press call
- Staffed live broadcast interviews
- Monitored the situation
- Got a full night's sleep before launch (!)
- Kept up with all the usual last-minute changes without the usual stress
If this is the AGI era, I am here for it! Excited for the world to get to experience GPT-6 Astra. It's a very good model and I think that it will transform both the way we work and how much we can accomplish.
Hiring Tony Robbins today costs 1 million dollars for just one day.
This is a 21-minute tape recorded in his own home more than 30 years ago.
There, he explains exactly how to get anyone to say yes.
The same material for which they now charge a fortune... completely free.
It's a rare, unfiltered recording from the time when he still didn't charge millionaires just for being in the same room.
When someone says no, they give two excuses:
“I don’t have time” or “I don’t have money”
Neither is true.
The real reason is that they still don’t believe it’s worth it.
It’s not a money problem. It’s a state problem.
Tony teaches something he calls “attack and confess”
Instead of arguing the objection, you confess your own:
“I had the chance to go six months ago and I didn’t until two months ago.
I can’t even imagine the time I wasted”
The room goes silent. No one argues.
Then he gets the person on the “yes train”
Each little yes adds to the next.
Until saying no at the end feels harder than saying yes.
When it’s time to sign, the person has already said yes five times without realizing it.
Just 21 minutes.
There you learn the two only moves that people pay 1 million a day for:
how to read anyone’s state… and how to shift it.
Most people spend years guessing in sales.
He wrote it on a flipchart in his living room in less than half an hour.
A seat in that room cost $125 dollars.
Today it costs $1 million dollars a day to sit in front of him.
The tape is free right now.
And the answer is in this video.
9 cool GPT 6 Astra prompts worth trying:
1. The bill renegotiator.
"Go through my internet, phone, and software bills, jump into each provider's chat support, and negotiate them down or cancel what I'm not using."
2. Turn an agency into software.
“Pick one service business in [niche] and reverse-engineer the exact workflow they sell to clients. Break it into steps, tools used, inputs, outputs, human judgment points, and places where the work gets slow or expensive. Then design the simplest AI product that could replace the first version of that service and charge $500-$5,000/month.”
3. Garage sale flipper.
"Watch Facebook Marketplace and Craigslist in my city for [cameras / furniture / bikes] listed way under market, and text me the second one's mispriced with the link."
4. Create my 1 person company dashboard.
“Look at my docs, notes, Stripe exports, analytics, customer calls, and project list, then build a weekly operator dashboard. I want to know what is making money, what is wasting time, what customers are asking for, what I should stop doing, and the three highest-leverage actions for next week. Be blunt and show your work.”
5. Audit my company for agent opportunities.
“Look at how this business works and find the tasks we should give to agents before hiring another person. For each task, estimate the current human time, the cost of mistakes, the tools involved, the difficulty of automating it, and the first safe version we could deploy. Prioritize things that save money or create revenue within 30 days.”
6. Be my browser operator.
“Use the browser to complete this workflow: [workflow]. As you go, click through the actual sites, collect the data, fill the forms where appropriate, and keep notes on what broke or slowed you down. When you’re done, give me the output, the repeatable SOP, and the automation plan so this can become an agent.”
7. The whole QA team.
"Every night, open my app on a real phone, go through signup, checkout, and the main flows, and screenshot anything that's broken or confusing."
8. The competitor spy.
"Sign up for my top 3 competitors, sit inside their product and their emails, and send me a monthly report on every new feature, price change, and thing they do better than us."
9. Make a game people would actually play for 5 minutes as a lead magnet
“Build a browser game around this mechanic: [mechanic]. Don’t just make a cute demo; add progression, tension, scoring, failure, polish, and one reason someone would send it to a friend. Then add lead capture (email/sms), it needs to tie into my core product which sells XYZ.”
A lot of people are asking how I pulled off these super long-horizon builds with Astra.
Astra is extremely powerful, but by default it struggled with a task this difficult. I tested a bunch of approaches to get past this, and the one I landed on is something I'm calling the Manager Loop.
It's basically a couple of tricks we used to use with much less capable models a couple of years ago, with a few new ideas layered on top. Turns out that when you put those together and apply them to Astra, its ability to do extremely difficult long-horizon tasks goes up dramatically.
Here's how it works:
1. Launch an agent (I'm calling this one the "manager"). Chat with it about what you want to get done, and have it build a massive checklist of to-dos, then break that checklist into phases.
2. The manager then spawns a second Codex agent in a separate thread (the "implementer"). The two agents can message each other.
3. Put the manager in /goal mode, and tell it to run each phase on the implementer in /goal mode.
4. The manager messages the implementer: "/goal Complete phase one completely, extremely well." The implementer doesn't stop until that phase is done, then messages the manager back. The manager tells it to start phase two. They repeat until every phase is finished, completely autonomously.
Why I think this works: over a long-horizon task, Astra tends to asymptote. It gets way further than previous models, but at a certain point it kind of just stops improving against the goal as quickly as it did before. It gets stuck in the minutiae, focusing way too much on small details, and overall progress stalls. The Manager Loop forces it to work piecemeal, one phase at a time. It's essentially how a human would steer a model, except the model is doing the steering for me.
That's actually how this started. I was having the model write the checklist and break it into phases, and then I was doing the manager's job by hand. At some point I thought, "Wait, why can't I just get a separate AI to do this?" That's what unlocked full autonomy, which is super useful.
A wording detail that seemed to matter: I ask for each phase to be done "extremely well," not "perfectly." Maybe I'm reading too much into it, but asking for "perfect" sent the model right back into the minutiae. "Extremely well" implies it's allowed to move on once it's good enough, and that worked better in my testing.
One more trick that I think helps (this one is more of a hunch, but it was useful for me): have the implementer build a simple HTML page with the full checklist on it. The implementer checks boxes off as it goes and updates a counter, and the page has a chart of # of boxes ticked over time.
Obviously the boxes aren't all equal, but it forces the model to notice things like "I haven't made progress in a while, time to move on." You can even put this in the prompt directly, like: "if you haven't ticked a box in X amount of time, move on". That helps a lot.
I also ran 96 sub-agents at a time. You can change this in your Codex config (or just ask Codex to change it).
This got me far better long-horizon performance than anything else I tried. I'll be sharing more in the coming days!
I occasionally get asked why I post so many visual things when new models come out. One reason is that nobody clicks any links so the visual stuff best communicates AI progress.
But if you want detailed reads, here is the research from my AI lab at Penn: https://t.co/hTOsEskT9H
Andrej Karpathy spent 8 years at OpenAI and Tesla
Last week he compressed everything he knows into one free 2-hour lecture
Agents → Loops → Harness → Graphs
People spend $15k on bootcamps that teach less than this
You probably don't have 2 hours right now
Don't let it vanish from your feed
Watch it, then read the graph engineering guide below
Today we release Pipette, a model evaluation suite for on-device intelligence, in partnership with @ArtificialAnlys
Most benchmarking platforms are optimized to measure core capabilities and speed profile of foundation models served in the cloud. Pipette gives the field a common, reproducible way to measure the quality, speed, latency, and memory use of AI models on devices such as phones, laptops, PCs, AI boxes, and embedded hardware.
> Pipette is open source
> In Pipette, models get compated as model + quantization + runtime + device from one interface.
> It comes with a warehouse of verified benchmark results, currently with 10k+ results across 35 model classes, 7 quants, llama.cpp runtimes, and 4 devices.
> Pipette is a dynamic platform, allowing new contributions from day one, adding new devices, runtimes, model families, and quantization levels.
🧵
Introducing Tenet, our first model post-trained for legal.
Tenet is a Kimi K3 base that we post-trained with @FireworksAI_HQ on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work.
Training increases Tenet's all-pass rate by 82% on LAB and 22% on LAB Contracts relative to the Kimi K3 base model. It achieves state-of-the-art performance on LAB Contracts and places second on LAB.
These gains generalize to other leading agentic benchmarks including @mercor's Apex Agents - Corporate Law, @crosbylegal's Redline Bench, and @scale_AI's Professional Reasoning Bench.
Tenet is also optimized for token efficiency, operating at less than a fourth the cost of leading foundation models.
We additionally post-trained three specialist models for Tenet to use as subagents:
1) M&A Diligence: post-trained with @baseten on our LAB Diligence environment in an RLM harness, this model is optimized for high-scale, long-horizon tasks.
2) Review Tables: trained with @appliedcompute on our Review Table environment, this model is state-of-the-art and cost-effective at high-volume document review and structured data extraction.
3) Firm Knowledge: trained with @EngramLab on our synthetic law firm environment, this model is optimized to learn and search over a firm's knowledge via memory and structured notes.
More details on model training, environment design, benchmarking, results, and more in the article by @gabepereyra below.
What's next for Harvey’s research?
- Scaling LAB to more jurisdictions, practice areas and workflows
- Scaling compute to bring new generalist models and capabilities to Harvey
More to come soon.
A Chinese developer just explained the shift from Loop Engineering to Graph Engineering better than anyone.
most people are still building agents the way that's about to be obsolete.
> why single-agent loops break and go "goal blind"
> the 4 parts of a graph: nodes, edges, state, policy
> 3 topologies that run everything: diamond, supervisor, pipeline
> Anthropic's 5 official workflow patterns
the punchline: it's not how many agents you run. it's the determinism you build with verifiers, code fallbacks, and reality anchors.
I broke the same architecture down with Kimi K3. Full A-Z guide below.
This went under the radar this week, but we just shipped the ability for models with multi agents v2 to delegate to any supported model, including Luna!
Took a bit of time to make sure this worked reliably