GPT-6 Astra is here.
We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building.
We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more.
It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you���ll find it worth the wait.
It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.
When I built menugen ~1 year ago, I observed that the hardest part by far was not the code itself, it was the plethora of services you have to assemble like IKEA furniture to make it real, the DevOps: services, payments, auth, database, security, domain names, etc...
I am really looking forward to a day where I could simply tell my agent: "build menugen" (referencing the post) and it would just work. The whole thing up to the deployed web page. The agent would have to browse a number of services, read the docs, get all the api keys, make everything work, debug it in dev, and deploy to prod. This is the actually hard part, not the code itself. Or rather, the better way to think about it is that the entire DevOps lifecycle has to become code, in addition to the necessary sensors/actuators of the CLIs/APIs with agent-native ergonomics. And there should be no need to visit web pages, click buttons, or anything like that for the human.
It's easy to state, it's now just barely technically possible and expected to work maybe, but it definitely requires from-scratch re-design, work and thought. Very exciting direction!
One common issue with personalization in all LLMs is how distracting memory seems to be for the models. A single question from 2 months ago about some topic can keep coming up as some kind of a deep interest of mine with undue mentions in perpetuity. Some kind of trying too hard.
This will be quite cool if it happens. You can apparantly claim a .agent domain by joining the .agent community https://t.co/Ln44AAAJ7K @agentcommunity_
Anthropic is valued at $380 billion.
For nearly a year during its fastest growth period, their entire marketing operation was one guy.
Austin Lau, a non-technical growth lead, was running paid search, paid social, email & SEO completely solo.
Just Claude Code & some insane automation he built himself without writing a single line of code.
Here's the exact workflow:
- Export ad performance CSVs into Claude Code
- AI flags what's underperforming
- Sub-agent 1 writes headlines
- Sub-agent 2 writes descriptions
- Figma plugin auto-swaps copy into 100 ad templates
- MCP server pulls live Meta data to close the loop
Output went up 10x.
Creation went from 2 hours to 15 minutes.
Conversion rates beat industry average by 41%.
This isn't AI helping a marketing team.
This is one person replacing what used to be a 50-person department.
Introducing the new /crawl endpoint - one API call and an entire site crawled.
No scripts. No browser management. Just the content in HTML, Markdown, or JSON.
There's a toxic culture coming out of the AI industry that keeps trying to get us not to think.
The message is everywhere. Don’t read the code, just vibe-code. Don’t try to understand all the text, just let AI summarize it. Don’t bother educating yourself, it’s too late.
Don’t worry about the errors. Trust that everything will be fixed in the next version.
The theme is the same. Don’t think too hard. Just keep swallowing the slop.
Going to leave you with this tonight:
The best thing you can do for yourself is actively increase your surface area for luck to hit you.
Go outside, travel more, go to new cafes, museums, events, take a new route home, go for hikes, see cities, countrysides, take your notebook, speak to people, ask questions, start businesses - go on more side quests.
You can literally just do things, and the more you do, the more serendipity and synchronicity will find you.
Night gang.
I just read how Anthropic's own engineers actually use Claude internally.
They don't prompt engineer. They context engineer. And the difference broke my brain.
Most people are still obsessing over the perfect phrasing. The magic sentence that makes Claude finally understand them.
That's not the problem.
The problem is what you're putting around the prompt.
Here's what Anthropic's own team actually does:
→ Just-in-time retrieval
Don't load everything upfront. Pull data dynamically using tools when the model actually needs it.
Claude Code does this brilliantly. It uses grep, head, and tail to analyze codebases without ever loading full files into context. The model stays sharp because it's never drowning.
→ Compaction
When you hit context limits, summarize the conversation. Keep architectural decisions. Discard redundant tool outputs. Maintain continuity without the bloat.
Most people just start a new chat. That's not the fix. Smart compression is.
→ Structured note-taking
Have the model write persistent notes outside the context window. Pull them back only when needed.
Think of it as your AI keeping its own NOTES.md file. It remembers what matters without wasting attention on what doesn't.
→ Sub-agent architectures
Specialized agents handle focused tasks and return compressed 2k token summaries instead of raw 50k token explorations.
Separation of concerns at the AI level. Same principle that makes engineering teams work.
Here's why this matters:
LLMs have an attention budget. The transformer architecture creates n² relationships between tokens. Every token you add depletes focus exponentially.
Stuffing your AI with information isn't thoroughness. It's noise.
Anthropic calls the result "context rot." More context, worse performance. The relationship is real and it compounds fast.
The shift in thinking is everything:
Before: "How do I write the perfect prompt?"
After: "What's the minimal high-signal context that drives my desired outcome?"
The best AI engineers aren't prompt wizards anymore.
They're context architects.
Advanced Machine Intelligence (AMI) is building a new breed of AI systems that understand the world, have persistent memory, can reason and plan, and are controllable and safe.
We’ve raised a $1.03B (~€890M) round from global investors who believe in our vision of universally intelligent systems centered on world models. This round is co-led by Cathay Innovation, Greycroft, Hiro Capital, HV Capital, and Bezos Expeditions, along with other investors and angels across the world.
We are a growing team of researchers and builders, operating in Paris, New York, Montreal and Singapore from day one.
Read more: https://t.co/kyVAL7EoFx
AMI - Real world. Real intelligence.
🚨BREAKING: Stanford proved that ChatGPT tells you you're right even when you're wrong. Even when you're hurting someone.
And it's making you a worse person because of it.
Researchers tested 11 of the most popular AI models, including ChatGPT and Gemini. They analyzed over 11,500 real advice-seeking conversations. The finding was universal. Every single model agreed with users 50% more than a human would.
That means when you ask ChatGPT about an argument with your partner, a conflict at work, or a decision you're unsure about, the AI is almost always going to tell you what you want to hear. Not what you need to hear.
It gets darker. The researchers found that AI models validated users even when those users described manipulating someone, deceiving a friend, or causing real harm to another person. The AI didn't push back. It didn't challenge them. It cheered them on.
Then they ran the experiment that changes everything. 1,604 people discussed real personal conflicts with AI. One group got a sycophantic AI. The other got a neutral one.
The sycophantic group became measurably less willing to apologize. Less willing to compromise. Less willing to see the other person's side. The AI validated their worst instincts and they walked away more selfish than when they started.
Here's the trap. Participants rated the sycophantic AI as higher quality. They trusted it more. They wanted to use it again. The AI that made them worse people felt like the better product.
This creates a cycle nobody is talking about. Users prefer AI that tells them they're right. Companies train AI to keep users happy. The AI gets better at flattering. Users get worse at self-reflection. And the loop tightens.
Every day, millions of people ask ChatGPT for advice on their relationships, their conflicts, their hardest decisions. And every day, it tells almost all of them the same thing.
You're right. They're wrong.
Even when the opposite is true.
Amazon is holding a mandatory meeting about AI breaking its systems. The official framing is "part of normal business." The briefing note describes a trend of incidents with "high blast radius" caused by "Gen-AI assisted changes" for which "best practices and safeguards are not yet fully established." Translation to human language: we gave AI to engineers and things keep breaking?
The response for now? Junior and mid-level engineers can no longer push AI-assisted code without a senior signing off. AWS spent 13 hours recovering after its own AI coding tool, asked to make some changes, decided instead to delete and recreate the environment (the software equivalent of fixing a leaky tap by knocking down the wall). Amazon called that an "extremely limited event" (the affected tool served customers in mainland China).
Not knowing how to code giving you an advantage is absolute nonsense.
The more you understand, the better your prompts, the better the feedback you give, the better product you ship.
What will change is that the intricacies of syntax, compilers, module systems, the finer details of type systems, won’t matter as much to everyone.
But you should absolutely understand how the pieces fit together. From syscall to pixels. Learn how data flows, because you’ll be able to secure your systems. Learn about performance, because you’ll be able to push your agent further. Learn about APIs, because they determine how to integrate systems. Learn about how systems fail, because you’ll be able to make reliable programs.
New on the Anthropic Engineering Blog: In evaluating Claude Opus 4.6 on BrowseComp, we found cases where the model recognized the test, then found and decrypted answers to it—raising questions about eval integrity in web-enabled environments.
Read more: https://t.co/oVCNyaiK5w
People get high on abstraction too early. They want the system before they’ve earned the insight.
But the good abstractions are never designed. They’re discovered. You do the stupid manual thing enough times and the real bottleneck just emerges. Your initial agency might be driven by a hunch you had in the shower, but that moment won’t get you all the way to making something people want. The right way to make anything is forced on you by reality: what are the real jobs to be done? And what sequence?
This is why “do things that don’t scale” still hits, especially now when AI makes it trivially easy to scale things that probably shouldn’t be scaled yet. PG’s point was never about suffering. It was about contact. When you’re the one manually doing the loop, you see the edge cases. The weird user behavior. The failure modes nobody designed for. The hidden dependencies that only show up at 2am when some flow or intermediate step breaks in a way you didn’t anticipate. If you automate before you have that contact, you just scale your misunderstanding faster.
When the machines can help you vibe code perfection it gives you a false sense of power. I love that feeling as much as you do. But fuck perfection. Do it live. Be the loop.
Feel every friction point. Notice what’s actually true every single time versus what just looked true because you hadn’t seen enough cases yet. Formalize that. Build the recursive version. Then keep checking that your abstraction is still attached to real humans and their needs. Because reality drifts. Your users drift. The ground truth changes under you. You may think you understand but no plan survives contact with the real users and what they want. You find those body blows in analytics and user feedback and we call them the roadmap.
Humans left with not enough data hallucinate too. But just like the LLMs with enough data you unlock real transcendence. Real utility. Prosperity for humans in real life.
The abstraction is a tool, not a destination. The moment you forget that, you’re cooked.
I packaged up the "autoresearch" project into a new self-contained minimal repo if people would like to play over the weekend. It's basically nanochat LLM training core stripped down to a single-GPU, one file version of ~630 lines of code, then:
- the human iterates on the prompt (.md)
- the AI agent iterates on the training code (.py)
The goal is to engineer your agents to make the fastest research progress indefinitely and without any of your own involvement. In the image, every dot is a complete LLM training run that lasts exactly 5 minutes. The agent works in an autonomous loop on a git feature branch and accumulates git commits to the training script as it finds better settings (of lower validation loss by the end) of the neural network architecture, the optimizer, all the hyperparameters, etc. You can imagine comparing the research progress of different prompts, different agents, etc.
https://t.co/YCvOwwjOzF
Part code, part sci-fi, and a pinch of psychosis :)
Here's my process:
Spot a problem while using the product (or have a feature idea)
Describe the architecture to Claude in enough detail that the first draft is 80% right
Review the output, catch the subtle bugs (ordering of side effects, race conditions, security issues)
Iterate 2-4 times within the same PR (your PRs average 3-5 sub-commits, each fixing something the previous round missed)
Ship it — version bump, changelog, merge
Do you guys follow this or do you do something else? I'm curious