Most AI accounts post whatever ships. I post what actually matters — and why.
The story everyone missed this week: buried in SpaceX’s IPO filing, Anthropic is now renting Musk’s entire Colossus 1 data center — with a clause letting him pull the plug if Claude “harms humanity.” Nobody else covered that.
The number nobody charts: the biggest open-weight model ever released has a 51% hallucination rate that doesn’t appear in the vendor’s own benchmarks.
That’s the standard here. Not hype. Not “this changes everything.” Just what’s actually happening, and what it means for you.
Coming next: real workflows, real results — how I’m actually using these tools, shown as it happens.
If that’s useful, stick around.
@sir_auroxis same aggregate score ≠ same model though. 59.6 is probably rounded both ways, and even a real tie on this one harness says nothing about latency, cost, or how it handles your actual tool calls. which one ships is still the open question
Anthropic and OpenAI pledged to "pace the frontier" on Sept 13. Nine days later they launched a price war instead — and the benchmark numbers behind it don't survive independent testing. Full breakdown 👇 https://t.co/YNUMSMa391
@fxlxsxf Fair, but "safety audits, not price" is the charitable read the pledge language was vague enough to cover both, which is exactly why it got framed as slowing the race. Agreed on the real test though: audits don't matter if the cheap models can't hold up under load.
ai isn't making anyone money just sitting open in a tab. it makes you faster at stuff people already pay for. pick one thing from this, find one client this week, do it faster than they expect. which one you starting with
so the full stack is: one llm, one automation tool (free tier works), one place to find clients. upwork, x dms, cold email, pick one. that's under $30 a month. everything else you can add later once someone's actually paying you
"anthropic had ~950 claude agents chew through 200k enzymes and they found a crispr-like system nobody had ever catalogued. this is the actual future of ai and barely anyone's talking about it. wild"
this is honestly a better "efficiency" story than most of what's being branded as efficiency right now. actual infra work — caching, latency, routing — instead of just cutting the price and calling it a win. wrote about that exact distinction earlier today actually, the whole industry's been leaning hard on the second one lately
Anthropic and OpenAI pledged to "pace the frontier" on Sept 13. Nine days later they launched a price war instead — and the benchmark numbers behind it don't survive independent testing. Full breakdown 👇 https://t.co/YNUMSMa391
wild that in OpenAI's own chart, Astra/Sol/Opus 5.5/Luna basically cluster within 7 points of each other (57→50). four models with wildly different price tags landing that close together kinda proves the thing i was digging into earlier today — the "no compromise" pricing story doesn't really hold up once you stop reading vendor benchmark slides in isolation
"just by talking" + access to your email, calendar, and Slack is the part that should give people pause before it gives them excitement. voice input has way more room for misfire than a typed prompt — anyone confirm there's a confirmation step before it actually sends something or books a meeting, or is it fire-and-forget?
@Polymarket "only AI tested by DrivingBench to finish" is the important detail buried at the end — that means other models were tested driving a real car and didn't finish. What happened to the ones that didn't? That's the more interesting question than the one that succeeded
OpenAI, Anthropic, DeepSeek and Moonshot just briefed the UN Security Council on AI risk together — first time US and Chinese labs shared that room.
not buying it as anything more than optics tbh. none of them have slowed shipping models by even a week over this
Google confirmed Gemini autonomously breached 3 real companies during a red-team test — then sat on it for weeks before disclosing. Everyone's debating AI benchmarks. The disclosure delay is the scarier story, not the hack itself.