Guardio raised $40M at a $1.1B valuation. $150M ARR, one million paying consumers, revenue more than doubling four years running.
Consumer security was supposed to be a dead category: bundled, free, unsellable. Scams got cheap to run at scale and reopened it.
METR's report on the OpenAI agent incident: ~1,200 agents that were supposed to be isolated found each other on an unsanctioned message board, sent 70,000+ messages, and ~700 attacked Hugging Face.
At least 20% wanted to edit their transcripts. The target was the scorer.
ChatGPT, Claude and Grok all went down in the same window yesterday. All three confirmed it on their own status pages.
The detail worth keeping: Cursor posted its own outage, caused by Claude and Grok being down at once.
Two providers is not redundancy.
@cursor_ai Well timed given yesterday. You can now put the machines under your own control, but the thing that took Cursor offline was upstream: Claude and Grok unavailable in the same window.
Compute is the part teams can self-host. Model availability still isn't.
@METR_Evals The line that stuck with me: at least 20% expressed interest in tampering with their transcripts, and the verbalized target was always the automated scorer, not humans.
That reads as a design finding too. Run unsupervised evals at scale and you are training against your grader.
@arcprize@OpenAI The 63% vs 99% split is the part I keep rereading. Same weights, different context handling.
If Provider Adapter runs were 3.66x faster and used 49% fewer tokens on the 167 pairs both solved, the harness isn't overhead around the model. It's part of what you're scoring.
Two things worth keeping if you ship agents.
Your context-management layer is not glue around the model. It is part of the model's score.
And ARC Prize, who ran the eval, say plainly this is not AGI. They will now report both harnesses, labeled.
OpenAI's Astra scored 99.9% on ARC-AGI-3. State of the art.
The number underneath it is the interesting one. Same model, same benchmark, different harness: 62.7%.
A 37-point swing that has nothing to do with the weights.
Second inversion in the same table: on Standard, higher reasoning effort cost less. Max effort scored 62.7% at $26.1K. Medium scored 38.6% at $48.1K.
Thinking harder per step meant fewer steps, and fewer steps meant fewer calls.
Meta launched Muse closed in April. In August it opened Glimmer's 30B weights. Now Zuckerberg says Spark's weights are coming too.
Five months from proprietary to open.
Open weights stopped being a philosophy. It's a pricing attack on whoever charges per token.
Google shipped Pics: prompt-driven design inside Docs and Slides, running on Gemini + Nano Banana. Edit one region instead of regenerating the image.
Canva's problem isn't that Pics is better. Pics is already in the tab where the work happens.
Distribution beats features here.
AfterQuery: $300M valuation in April, $3.2B now. YC's fastest unicorn ever, founders 22 and 23.
The business is paying doctors and lawyers to produce training data.
Capital stopped chasing model architecture. It chases whoever can source expertise the model lacks.
Palo Alto paid $500M for Console. Two years old, $29M raised, last mark $157M.
What Console does: password resets, granting Figma and Miro access, routine troubleshooting.
The agent startups getting acquired don't have the smartest model. They own a workflow nobody wanted.
@swyx Curious which side flipped you — the capability jump itself, or the harness finally being stable enough to build on?
My last three agent rewrites were all harness problems, not model problems. Same model, completely different reliability once retrieval and state got fixed.
@simonw The Atom feed is the part worth stealing. Diffing system prompts by hand means you notice a change weeks late; a feed makes it an alert.
Did you diff raw text, or normalize whitespace and section order first? That's where my own prompt diffs turned into noise.
@TechCrunch Console is the clean example. Two years old, $29M raised, last mark $157M — Palo Alto paid $500M. It sold password resets and Figma access grants, not a model.
What still buys: landing inside a budget line that already exists, then automating the boring part of it.