I really hope everyone copies the work we did here. I’ve been thinking about how many tokens are wasted with bad PR monitoring flows and it made me feel a bit sick
aha! with the benefit of hindsight, i just figured out openai’s marketing team seriously cooked hard with their gpt 5.6 release
let me reverse GTM what happened for you
at that time, fable was already released and people were loving it. but openai knew astra was far from being ready
if people ask “what’s your answer to fable” and they say “we are a few months behind”, it’ll damage their brand
so instead of doing the typical 2-tier “gpt 5.6” and “gpt 5.6 mini”, they artificially created a 3-tier system with a new naming scheme - sol, terra, luna
pretty sure terra was not even originally planned but added last minute to make this plan work. in their release announcement, terra was said to be the same performance as gpt 5.5, while sol is a step function better
this created a perception that sol was openai’s fable equivalent. if you still remember, most of the reviews back then were comparing sol with fable. many of which even said sol was better
that bought so much time for openai to catch up without being seen as falling behind
and now that astra is here, terra immediately got killed in the 6 series because it was never real!
everything made sense now, and it’s absolutely brilliant
Max velocity. All new code:
No tests or automated validation; bare-minimum checks.
Ugly is fine if it works.
Reuse, don't reinvent.
Verify hands-on, not by tests or looks.
Fix every bug, re-verify, no regression tests.
Test blocking you? Delete it for good, then fix the cause.
@Rosemarry__vt RA II or PvZ? I grew up watching my dad play them. I still play, but sometimes it feels more like keeping up the identity of “someone who loves games.” I’m not tired of gaming, I just don’t think I’ve loved anything like I loved PvZ back then, not even PvZ now.
The worst part about AI engineering is how often does everyone think their weird esoteric bullshit actually matters.
I think it's a combination of sycophancy + non-determinism + models getting better so fast that you can associate model improvements with random changes you made to your environment. The result is that everyone's brains fall out as they overthink things that don't matter.
almost any frontier model would get this immediately. astra somehow comes up with 3 wrong explanations and then spends time testing them. genuinely weird behavior.
almost any frontier model would get this immediately. astra somehow comes up with 3 wrong explanations and then spends time testing them. genuinely weird behavior.
Huge thanks to code002-2 and the upstream contributors for bringing Linux to sheng. Their kernel, device support and archlinux-sheng base made this possible. Our work builds on that foundation to bring Omarchy to the tablet. 🙏
https://t.co/2ANBlVPd5i
Omarchy on sheng! 🚀
Got @dhh's Omarchy running natively on the Xiaomi Pad 6S Pro 12.4. ARM64, Hyprland, 144Hz—and the keyboard, trackpad & touchscreen work.
My first Omarchy experience, on a tablet. Thanks, Astra, for helping make it happen!
Astra is rolling out now! It's unbelievably powerful, but you have to push it a bit to really see the difference.
I wanted to give some examples of what I mean here. Here are some fun things to try in your codebases when you get access:
1. Slop audits
I have asked Astra to go through all my codebases hunting for slop. Useless tests, unnecessary function wrappers, stuff like that. Surprisingly effective. It's cleaned up a ton of code for me.
2. Hunt for performance wins
Astra has found a ton of performance wins in my apps. It can make hard cuts and verify the results. Make sure you give it the tools it needs to verify it's changes. On that note...
3. Improve agent dx and verification loops
This model is surprisingly aware of what it can do. Ask it what it needs to verify it's own work. Let it suggest improvements to setup and worktree flows, debug access, end to end QA flows, etc.
4. Audit open PRs/issues
I've had Astra close at least 200 PRs and issues at this point. It also does a great job of finding easy win PRs to merge. Super helpful for projects with lots of contributors (both OSS and internal repos)
5. Let it merge
Once you get used to the model and it's ability to verify work, try trusting it a bit more. Obviously don't let it yolo ship to production, but if you have a good flow for pr -> main -> staging -> prod, maybe roll the dice a bit?
6. "takeover" work that is stuck in a loop
I've had this model land PRs that were stuck for months. If you have some old branch or PR where your agents are running in circles, blowing up the spec without actually shipping, tell Astra to take it over. Make sure it knows it can throw away the existing work and start from scratch.
Hope these help you guys with really pushing the new model! I've had a blast with it, I hope y'all do as well :)
Note: I cannot be held responsible for surprise bills and limit usage
🤔 Codex compaction feels like “magic,” but it can still lose details from the original task.
What if, before compaction, the full chat history were turned into an index the agent could reference on demand?
Could this improve task consistency across compaction?
I'v extract gpui-base from gpui-component. It provides base GPUI controls and interactions without any style, making it easy to build UIs with different designs.
gpui-component is now built on top of gpui-base, with a default style ready to use.
https://t.co/6jLlodfJlH
opus 5 is a sign that the obsession with “long-horizon agents” in model training is finally backfiring
i don’t like long-horizon agents, and i’ll explain why they fundamentally don’t work
some people will immediately jump out and say “skill issue”. well, show me one profitable business you built with a long-horizon agent working all by itself - i’d love to learn
so far, the only thing they were able to build that’s even interesting enough for people to talk about are those 3d games that are a partial clone of something that already existed
the reason an agent was able to build a working prototype of complex games like call of duty was that a team of humans already figured out all the requirements years ago for how such games should work, what kind of controls are intuitive, what mechanics are fun etc
all those requirements were already absorbed into the model weights, so when you say “build me call of duty” the model already knows the details. its long horizon execution capability can get all the requirements implemented, which i must say is indeed impressive
but now you can see - the value of long horizon execution has a prerequisite of a massive amount of high quality requirements clearly defined upfront. it took a big team of very talented humans months of effort and many iterations to define that for call of duty
now imagine games like call of duty don’t exist yet, how would we use agents to build it for the first time? we can’t say “build me call of duty” any more. and there’s no way we can define months-worth of game design details upfront
we’ll have to build a tiny prototype of the most basic mechanics, play with it, see if it’s fun, then iterate and expand the complexity. even with the smartest humans, that’s how we work towards something great
we don’t need agents to go dark for a long time, spend tens of thousands of dollars worth of tokens, and come back with a product the agent randomly decided to build - try build something truly novel with this and you’ll see it can’t come up with anything that’s actually profitable (i’ll show you why in a bit)
we need a tight feedback loop where we can collaborate with the agent, plan with it, understand what it’s done, question its approach, apply our judgement, give it real world feedback and iteratively arrive at a good outcome
and that’s exactly what opus 5 absolutely suck at. why? i explained it in more depth with my previous post on how RLVR works - RLVR trains the model to generate code that can pass predefined tests in an isolated environment, which is fundamentally incompatible with the idea of having human in the loop. the more we train the models with RLVR to be “long-horizon”, the less they care about talking to humans
ok now - why do they have to talk to humans? why can’t the models iterate and apply judgement by itself?
maybe one day they could, but not today, due to many limitations. two examples -
1. LLMs today can’t “watch a video” yet. they can look through a lot of screenshots, which is extremely inefficient at observing a high fps animated signal. so anything that requires continuous visual attention is something LLMs can’t do very well
2. LLMs don’t truly understand what’s “intuitive” or “pleasant” for humans. they know what’s already proven to be intuitive and pleasant in the past, but if you present a truly novel concept, it can’t predict whether humans will like it accurately
because of those limitations, human judgment is still needed for almost anything valuable. without humans in the loop, agents will only be able to repeat something that already existed, or go in random directions without true understanding of whether it’s building something useful
in summary, long horizon agents assume requirements all exist upfront. they are fundamentally against human in the loop. and they don’t have true judgement for what humans like
that, my friend, is why they don’t work