ok it's official... AI-written unit tests and integration tests are empirically proven to be unhelpful. you should tell your agents to stop writing tests by themselves
on the deepswe eval set, banning sonnet 5.5 high from writing any tests actually resulted in slightly higher success rate (non stat-sig), with less time and token spent (stat-sig). this is pretty hard empirical evidence
across the tests that were written by the "tests allowed" baseline arm, 65% of them were unit tests, 35% were integration tests, neither bucket resulted in any improvement compared to not writing any tests at all
i also picked a random sample of 44 tasks subset where i completely disabled executing even existing tests - it also did not affect success rate at all
i spot checked many tests written in the baseline arm, and my intuition is that most tests are simply a repetition of the implementation
agent-written tests do not add any value because both the implementation and the tests were simply the agent's interpretation of our intent. the tests aren't any more accurate than the implementation itself
i suspect we can still extract some value from unit tests and integration tests if we describe them ourselves when we believe we can articulate our intent better through test cases than through requirements. but i have not proven this yet
also worth noting, during deepswe eval the agent wrote almost no e2e tests (only 17, compared to 3000+ unit/integration tests written). so this analysis does not prove nor disprove the value of e2e tests - i will do another eval specifically for that
.@TravelGov: Irkutsk, Rusia: La Embajada de Estados Unidos está al tanto de los informes de prensa procedentes de la región de Irkutsk sobre un presunto caso de peste neumónica que resultó en el fallecimiento de una persona, así como de medidas de cuarentena y el cierre de hospitales en Irkutsk. Seguimos monitoreando la situación de cerca. El gobierno de Estados Unidos tiene una capacidad limitada para asistir a los ciudadanos estadounidenses en Rusia, especialmente fuera de Moscú. La alerta de viaje de Nivel 4 del Departamento de Estado («No viajar») para Rusia sigue vigente e indica que los ciudadanos estadounidenses que se encuentren actualmente en el país deben abandonarlo de inmediato. Se recomienda encarecidamente a los ciudadanos estadounidenses que planeen viajar a Rusia que no lo hagan, debido al riesgo de detención arbitraria y a preocupaciones de seguridad.
A lab worker in Russia broke a test tube full of plague. Three days later she was dead.
Russia quarantined 189 people, locked down the lab, and then announced her autopsy showed no plague at all.
Those two facts don't sit together. You don't quarantine 189 people and hand plague-exposure questionnaires to pregnant women at a maternity center over pneumonia of unknown cause. Russia's chief sanitary doctor flew 2,600 miles to Irkutsk to discuss what local officials would only call "a particularly dangerous infection." The governor next door wrote that the woman died of plague, then edited his post to say "possibly."
Watch what they quarantine, not what they say.
The deeper story is why a plague institute exists in Siberia at all. Plague never went away. Yersinia pestis lives permanently in marmots and ground squirrels across the Mongolian border, and Buryatia reported an outbreak near that border this year. The Soviet Union built a network of anti-plague institutes a century ago to monitor those reservoirs, and Irkutsk was one of the big five.
That network led a double life. It did real public health work, and it also supplied virulent strains to the Soviet bioweapons program and trained its scientists to handle them. Igor Domaradsky, one of the founders of Biopreparat, ran the Irkutsk institute before taking that job. The five Russian anti-plague institutes are still closed to outsiders today.
Here's the part that should lower your heart rate. Pneumonic plague is the opposite of COVID as a pandemic threat. It incubates in 1 to 3 days, makes you visibly and violently sick, and kills untreated patients too fast for them to travel far. There is no silent spread. Generic antibiotics cure it if treatment starts within about a day of symptoms.
So the biology does the containment for you in every scenario except one. The scenario where officials spend the antibiotic window debating whether to say the word.
The pathogen is from the 14th century. The cure is from the 1940s. The only modern variable is how fast a government admits what it's looking at.
Prompt of the day:
/retro read my last 10 coding agent sessions and find ways to make my repo easier to navigate. Find where agents take too long to find relevant information, or rely on out-of-date docs.
Improving navigability is such an underrated way to save tokens.
people who laugh and say karpathy's latest tips are outdated - i suspect the vast majority of them have not even tried the tips yet
i just tested the ASD-STE100 wording rule and it's surprisingly good at helping increase clarity of model response, even within html artifacts. but the trick is that the full ruleset is a bit too strict and you need to pick a subset
one prompt you can run super easily:
"randomly sample 10 session transcripts where i worked with you interactively within the past week. apply ASD-STE100 rules to assistant responses and analyze which rules would have increased clarity, reduced confusion and improved the conversations, then document those rules in my user level AGENTS.md"
you might be impressed!
Google Mantis is a skills pack for security review with coding agents
Install:
npx skills add google/mantis
Key commands:
/mantis-threat-model: builds a threat model from your codebase
/mantis-researcher: scans for vulnerabilities
/mantis-review: filters false positives
/mantis-reproduce: writes a PoC and runs it in a sandbox
/mantis-patch: applies a fix and confirms it blocks the PoC
/mantis-report: generates the final security report
This is a good use case for running your agent in a sandbox, since the reproduce and patch steps execute generated code
https://t.co/OhmZDDNcWy
this has become one of my most used prompts recently:
> restate in your own words what you think my goals are and what the problem i'm trying to solve is
Last year, Anthropic was optimizing inference across 3 different types of compute (Nvidia, AWS trainium, Google tpus). There were mistakes made during these optimizations that resulted in actual degradation. Initially Anthropic denied it, but eventually they found, fixed and explained what happened. There have been no notable instances of degradation since.
Sadly, this one instance has made half of tech twitter’s brains fall out.
Separately, there’s the statistics side. You’re more likely to see the stupid spikes over longer windows.
If a model has a 1/50 chance of doing weird shit, and you do 20 prompts a day, there’s a ~30% chance you’ll have encountered weird shit on day 1.
By day 5, it’s closer to a 90% chance 🙃
Tl;dr, people are stupid, don’t understand non-determinism, and it happened once so they feel righteous
Okay, my Opus 5.5 benchmark with @VulcanBench Frontier v4 is in, took four days to benchmark this one, but was worth it.
Super interesting results. Opus 5.5 actually performed better at both Medium and High than at Extra High and Max.
And holy moly does the cost increase at Max, but you're not getting more accuracy.
While Fable 5.1 comes in as the most accurate, it's only at Max, but surprisingly, at a lower cost than at Extra High, which ties with Opus 5.5.
Also, important to remember, this is my Frontier v4 benchmark, these are super hard tasks designed to stump frontier models, not representative of what you'd actually give to these models on a normal day.
So my takeaway is, Opus 5.5 Medium is the way to go for a very strong balance of high accuracy and solid cost efficiency.
I see no reason to use Fable 5.1 ever, and Astra, while a solid model, just isn't on the same level at Opus 5.5.
So I think I can now confidently say, with data to back it up, that at this moment in time, Anthropic has officially taken the lead.
Will be adding all the details of this benchmark for anyone who wants to do a deep dive into the data to the VulcanBench site soon, stay-tuned 🖖
alright! as the dust settles around this insane week of model releases, i've stabilized around a new model line-up so sharing here for reference
this time my approach is a bit more structured. i've bucketed various LLMs into a few categories:
1. interactive orchestrators
these are models that i use as firstmate (and second mates), because they are fast, efficient, pleasant to talk to, intelligent enough to understand my intent, and have good enough judgment to steer the crew around it
for me, this bucket is opus 5.5 and grok 4.7
viable budget alternatives when my subscription quota runs out: muse spark 1.3, deepseek v4 flash
2. premium intelligence
these are models that i only use for highly ambiguous or creative tasks that i decide to truly need the extra intelligence and justifies the cost
i also use this bucket to handle escalations - when crewmates started arguing with each other, when a review-loop started spiraling out of control, when a simple change somehow ended up with a giant PR - i call these models to untangle the mess
for me, this is currently gpt 6 astra and fable 5.1 (or opus 5.5 when fable quota is tight)
3. planners
these models are the ones i trust as default for planning new features and investigating complex bugs. they would produce a spec that get implemented by a cheaper model
i almost exclusively use opus 5.5 for this right now because of its incredible ROI, whenever i don't need the premium intelligence
4. implementers
these are efficient workhorses that when given a well defined spec they can produce solid implementation
i use opus 5.5, gpt 6 sol, and grok 4.7 for this right now, and again the budget options: muse spark 1.3 and deepseek v4 flash (i'm sure there are many other viable alternatives as well - i just haven't got enough time to try them out)
among these models i currently find sol to be the best adversarial code reviewer (i haven't tried astra for this, as it's a bit too expensive to run at such high volume)
5. trivial fixers
these are the models i use for extremely trivial changes like a one-liner fix or config change. they really don't need much intelligence because often times what needs to be changed was already defined
for these i use gpt 6 luna and again the budget options if quota is tight
the way i actually make use of the categorization is that i told these routing preferences to firstmate, and it can then help me route the right task to the right model (which is now made very efficient because of Jev)
one last thing i'll point out is you can see how versatile opus 5.5 is in my line up - it can do pretty much everything! this is the first model that spanned across almost every bucket in my setup, which is quite a massive advantage because it means if i have to choose only one subscription it would have to be anthropic at the moment