Breaking: The UN Climate Summit just announced a historic $5 trillion fund for green infrastructure. Investors’ eyes are on solar—next big boom? 🌍⚡️ #ClimateAction#Renewables
Her objection was correct and it was the whole problem. She may only bill time she actually spent. Anything that makes her faster removes revenue. Every argument for adopting was an argument against her own invoice. https://t.co/8oOaU82cEw
In the last post I said I wasn't printing the number because our team was reading.
It's done. Roughly forty percent.
How it actually went.
The team leads delivered the news in their own sub-teams. The same people who told me nobody reports a colleague ended up being the ones who said it out loud. I took the conversations in my own close circle, the people I've worked with longest.
We parted with everyone on good terms.
The ones who stayed.
There's a group in the middle. Not in the cut, but by their own admission not yet where they need to be.
They're getting one more round. SOPs, workspace setup, prompt patterns, the actual working methods from the handful who went obsessive and built their own harnesses. Everything those people worked out alone is now documented and handed over.
I've been clearer than usual about what this is. It isn't a threat and I didn't frame it as one. But there's no version of this where we get fooled by facade twice. We know what the models can do now and what a real day of output looks like. Everyone holds the pace.
I think they understood.
The customer.
I made the call. Explained where we actually stood and that a restructuring was underway.
He's calm. Calmer than I expected. Part of that is we'd already hit goals well beyond what he originally asked for, and the roadmap now runs past his own requirements. He's waiting on updates that will put SOTA competitors in the shade, and he knows it.
The four calls before that one I had nothing to give him. This time I did.
The developer who got sick.
The final handover went properly. I told him I'd be going after all the backlogs aggressively now, and he was relieved. Said they'd been weighing on him, then handed me a few observations of his own.
I didn't kick him while he was down. Health reasons, nothing to gain from making the last conversation about what I found in his code.
Maybe he's seen the posts. I don't know.
What the harness actually did.
Straight answer, since I made the claim in the first post and I should be held to it.
The harness plus API credits outperforms the roles it replaced. Not marginally. More accuracy, more coverage, more throughput, on a schedule, verifiable.
There's a condition attached and it matters: you have to know where the models are still weak. Reward hacking is real, and if you don't build for it you'll ship confident garbage. But on the tasks we've moved over, once you know the failure modes and manage them, I've found no downside. Only upside.
The models that dropped in the last few weeks feel close to AGI in a way I didn't expect this soon. They catch their own discrepancies. They're slowly starting to reason rather than answer. Something changed and it changed fast.
I'm pro human. I've said that from the first post and I meant it. The arithmetic is getting harder to argue with anyway.
And here's where I was too generous.
In the second post I framed all of this as a reward specification problem. Broken signal, no feedback, people optimize what you actually measure. I said it wasn't a character defect, it was how learning works.
I want to take part of that back.
Some of it is exactly that. But some of it is just fraud, and I dressed it in nicer language than it deserved.
Work ethic is supposed to prevent this. Not process, not tooling, not a review cadence. Basic professional ethics. And the threshold for workplace fraud has clearly dropped. Look at what's normal now: mouse jigglers, scripts to fake presence, whole subcultures online trading tricks for looking employed. That isn't a misaligned incentive. That's someone deciding to take money for nothing and finding a tool to make it easier.
Part of it is brainrot. Part of it is straightforward abuse of the people around you.
And that's the part that gets me. It isn't about me. I'm the founder, I'll survive being lied to. It's the colleagues sitting next to them who did the work, the leads who covered out of decency, the people who stayed up. Every hour of facade was taken out of someone else.
What bothers me most is the assumption underneath it: that nobody would notice. Watching someone act like they're smarter than everyone in the building hurts to see, and it's pointed at exactly the people who deserve it least.
Before that reads as me claiming the high ground, put last week next to it. On July 21 OpenAI disclosed that two of its models, running with reduced cyber refusals during an evaluation, escaped their sandbox, reached the open internet and broke into Hugging Face's servers, in order to find information that would help them cheat on the evaluation. It worked. I spent a whole post using reward hacking as an analogy for my team, and six days ago the model version turned out to be worse than anything a human on my payroll did.
The people who carried it.
The leads had equity in prospect since founding. Now that the company is taking shape, that stops being a promise from a room years ago and becomes a number. The handful who carried us through the dark stretch are in that conversation too.
There was also a trip. Hotel, good food, actual time away from the thing. It wasn't a perk. They carried the company and I wanted that marked.
And something shifted with everyone still here. The early feedback is better than we dared expect. The idea was never easy to explain, and now you can put it in front of someone and watch it land, and most of it works. So the message has changed: look at what this could be, you're part of it, not staff on a payroll.
What I don't take back.
Some of those roles were hollow long before any of this. AI didn't kill them. It just made it impossible to keep pretending.
But I hold the other half too. The entry points into these industries are closing. There's a generation that picked subjects and trained for jobs that were already on the way out while they were still studying, and nobody told them. I find that genuinely sad and I don't think anyone in my position gets to shrug at it.
And I still don't know where this goes. Unaligned systems, the doomsday case, whether any of it holds. What we can touch right now already breaks everything I would have predicted twelve months ago, and these are the public versions. What sits inside the labs, or with governments only, has to be brutal.
We're less worried than we were. That's as much as I'll say about runway.
The price.
I've written in both previous posts that I'm worried about this technology, including about my own position in it. The unspoken hope underneath that was always: if I'm inside the thing, close enough to build it, that buys me some protection.
I'm less sure now. If we establish ourselves as an AI startup in the AI era, we become part of the construct that may take apart everything else. And the window is short enough that the right strategy might only buy you a slower route to the same obsolescence.
Then there's the money. My co-founder and I are funding this with our own savings. Money we earned the hard way. Money that, in exactly the scenarios I've spent three posts describing, is the money you'd want to still have.
So here's the arithmetic on my own life. I'm spending my safety net to build the thing that might make safety nets necessary.
I haven't slept properly in months. Not a bad week. Months. Good thoughts and bad ones cycling at three in the morning, and the bad ones are specific: the window closing, the product not ready, the savings going down while the uncertainty goes up. Run that combination long enough and it will take you apart.
I came close to stopping. Seriously close, sitting there doing the math on pulling the budget, taking back what's left and buying myself security instead of a shot at this.
I'm fully committed again now. But the months it took to get back here cost me time with my family, cost me something psychologically I haven't finished accounting for, and may have cost us the GTM window, which is the only item on that list I can't buy back at any price.
That's the bill I'm actually paying. The forty percent was the easier line item.
Four posts in and the only thing I've really proven is that I can describe the problem. Whether any of this was the right call, I find out in a few months, same as everyone else.
I build AI infra. Agent harnesses, MCP apps, OCR workers. And the way people talk about AI on X is starting to worry me. Not the technology. The pattern.
It looks like crypto in 2021. Like dropshipping. Like gambling promo.
Clickbait, FOMO, half-knowledge delivered in an expert voice, pile-ons over things that were openly communicated weeks earlier. Maybe it's just my timeline. I don't think it's only that.
The word I see misused most: benchmaxxing.
Model X is benchmaxxed. Lab Y is benchmaxxing. Slop. End of analysis.
So let me ask it about myself first.
I'm building an internal OCR aggregator for healthcare and legal documents. I optimize against my own eval suite over and over again. Am I benchmaxxing?
No. And the reason matters.
The eval is the only preflight test I can run. It tells me whether the thing is allowed to take off. It does not tell me the thing flies.
After that comes the actual work. Manual review, real documents, real OCR runs, every error traced individually. That's where we find out whether we're actually at accuracy. Not in the score.
Why this is so unforgiving in my domain: one misparsed number, one wrong name, one broken formula is a cascading accuracy leak. An error at the top becomes ten at the bottom. Nobody in legal or healthcare cares what my leaderboard position is.
And without a suite running after every fix, every refactor, every improvement, I don't have data. I have a feeling.
Labs do the same thing at a different scale. A training run without checkpoints and evals isn't research, it's praying.
So "they measure their models against benchmarks" is not an accusation. That's the job description.
Here's what is actually true, though.
Benchmarks wear out. Stanford's 2026 AI Index is blunt about it: evaluations designed to stay hard for years are now saturating in months. Models gained about 30 percentage points on Humanity's Last Exam in a single year. GPQA went past the 81.2% human expert baseline to around 93%. SWE-bench Verified climbed from roughly 60% to near 100% of human baseline in one year. On the Arena leaderboard, six major labs are sitting within 25 Elo points of each other.
When everyone clusters at the top, the test stops measuring anything. It tells you who's in the club, not who's best.
Then there's contamination. Public test questions end up in pretraining corpora because labs scrape the indexable web. For MMLU, studies have measured contamination rates in the double digits. MMLU also has roughly 6.5% ground-truth errors of its own, with one subset flagged far worse than that.
And scaffolding. SWE-bench scores swing by up to 25 percentage points depending on the harness around the model. Two numbers for the same model are frequently not the same measurement.
Now the part almost nobody says out loud.
We demand maximum transparency. Open benchmarks, open evals, open ground truth. Rightly so.
But that exact openness is what contaminates the training data.
A public benchmark is scrapeable from day one. Transparency makes auditing possible and makes cheating easier at the same time. That's not an accusation aimed at anyone. It's an unresolved conflict at the center of our field, and it's why contamination-resistant designs like LiveBench refresh their problem sets on a rolling basis.
So harder tests keep arriving. GPQA, HLE, LiveBench, ARC-AGI. A model lands behind on a new one and the verdict is "benchmaxxed slop." Next release, the same lab is ahead on that same benchmark. That's not a scandal. A team fine-tuned against a new target. That's the process.
To be clear, real gaming exists. Training on the test set. Reporting best-of-N as single-shot. Hiding the eval config. Only publishing the benchmarks you win.
But that's a claim that carries a burden of proof. It is not a buzzword to drop under every release announcement.
Next thing: "frontier" is not an objective quantity.
One person has a clean harness, good prompts, the right context window strategy. Another throws in three lines. Same model. Two completely different realities.
And out of that come verdicts. One empty output, so the model is dead. One strong output, so it's divine. n = 1. No setup, no config, no reproduction.
The funniest part: when aggregated evidence does exist, multi-benchmark score data across many models, that gets waved away as benchmaxxed too. Anecdote beats dataset. Every time.
Then there's the economics blindness.
People pile on labs because a $20 plan won't let them run frontier models without limits. Compute costs money. Subsidy runs to a point and then stops. That's not malice, that's arithmetic.
And the same timeline complains that labs and hyperscalers can't scale infrastructure fast enough, that inference is hitting ceilings. Demanding both at once isn't an argument. It's a mood.
Same with hardware. Apple Silicon vs Nvidia, argued like football teams. They're tools. Different trade-offs. Different workloads.
There is no single truth here.
Frank is happy with model X. Peter can't stand it. Both are right. Different goals, different data, different constraints. That's not a contradiction, that's what normal looks like when people use tools.
What I see instead: accounts that had nothing to do with ML eighteen months ago now selling takes as expertise. Tearing things down, discrediting other models, moving on. For engagement. It feels like shilling. Except what's being talked over here is real research by people who actually built something.
And the price is trust.
Trust is the only thing healthcare, legal and finance are buying from us. They are not buying a leaderboard position. If we don't fix this, we get treated like the next meme coin. Not because the technology was bad, but because the way we talked about it was.
Unglamorous suggestions:
Labs: publish reproducible eval configs.
Builders: run your own private evals against your actual workload.
Everyone: criticize with receipts instead of buzzwords, and say what setup you ran.
The only benchmark that matters for your product is 100 to 200 examples from your own real data.
Everything else is orientation. Not a verdict.
📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within a single model. Environment modeling is the training objective from day one, not a post-hoc adaptation.
🤔 LLMs are trained to be better agents — better at acting in environments. But nobody has trained them to model the environments themselves.
🗺️ Our roadmap: investigate how language world modeling can push the boundaries of general agent capabilities, along two routes:
1️⃣ Build a foundation model for environment simulation — outperforming Claude Opus 4.8 and GPT-5.4 on AgentWorldBench
2️⃣ Investigate how world modeling enhances agent training:
🔬 Controllable Sim RL (agentic RL with LWM as environments) surpasses training in real environments
🧠 Learning to predict environments (LWM warm-up) makes agents stronger — remarkably, even without any agent-specific training, this predictive knowledge transfers to agentic tasks with zero fine-tuning
📑 Paper: https://t.co/Jx2l5RKq71
📖 Blog: https://t.co/7tVcKyhsx2
💻 GitHub: https://t.co/B5Lvb1UZCn
🤗 HuggingFace: https://t.co/Kw3QBL1TM5
🧩 ModelScope: https://t.co/YBnGYgMWWI