Rumor mill says Claude Opus 5 launches today. No official word from Anthropic.
What’s real: a Polymarket contract on this exact question, trading at 29% for a Thursday launch, 91% for sometime this month.
The model is speculative. The bet on the model isn’t.
Claude Code just started grading its own homework before you even ask.
There's a new security plugin, live in beta. Every edit gets a free pattern check, zero extra cost.
Every turn gets read by a separate Claude with fresh eyes and no attachment to the code it's judging.
Commit or push, and a deeper agent goes through the whole diff like it's trying to break it.
Same inference you're already running. It just started working a second job.
Install it from the plugin marketplace and see what it's been catching without telling you.
I wrote an article about verifying which exact weather station a Polymarket market resolves on. Someone else skipped that step. They hacked the station instead.
The assumption everyone trading weather markets makes: the station data is ground truth, untouchable, the one thing you don’t have to question.
It isn’t.
→ Météo-France filed a complaint after one of its digital weather probes started reporting tampered readings
→ investigators found the sensor had been deliberately altered
→ the fake readings fed straight into Polymarket’s weather markets, triggering payouts on bets rigged in advance
→ Paris cybercrime prosecutors opened an investigation in May
→ ten weeks later, France ordered ISPs to block Polymarket outright, days before the World Cup final settled over $4B in bets on the platform
My article was about reading the referee correctly. This is what happens when someone stops reading the referee and starts controlling it instead.
Different threat model entirely. Verifying the station protects you from being wrong.
It does nothing if the station itself is lying.
The World Cup opened with a draw and closed with one too.
Spain started their campaign 0-0 against Cape Verde, group stage, matchday 1.
That exact scoreline was the biggest single hit on my flat $100 -on-the-draw thesis, +$1,300 on one match.
The final: Spain vs Argentina. Level 0-0 after 90 minutes.
Argentina became the first team in World Cup history to finish a final without a single shot on target in regulation.
The draw bet was live right up until Ferran Torres scored in extra time.
Tournament-wide: 7 scoreless draws, tied for the most in any World Cup ever. 9 of the first 24 group matches, over a third, ended level.
Started this thread at +$2,328 after 28 matches. The pattern never broke, the whole tournament through.
Betting the “boring” outcome was never boring this year.
+$2,328 edge at the World Cup - from betting the outcomes nobody wants.
You could literally automate the whole thing:
→ pull every match as it finishes
→ let Claude price the draw + “both teams score” markets
→ fire the bets the crowd refuses to make on Polymarket
Crowds misprice the boring outcomes.
Agents don’t get bored.
That’s the entire edge
Two weeks of silence. I wasn’t gone. I was building the thing instead of tweeting about building it.
Here’s what ate two weeks.
THE SETUP
Polymarket runs daily temperature markets. Highest temp in London, Istanbul, NYC, dozens of cities, every single day.
Most traders make one of two mistakes:
→ check a weather app, buy whatever number it shows
→ follow whatever bucket is already priced highest
Both lose money over time. Not because forecasts are bad. Because of one detail almost nobody checks.
MISTAKE EVERYONE MAKES
These markets don’t resolve on “the city.” They resolve on one specific weather station, usually an airport sensor, buried in the rules text.
City center and airport can differ by several degrees on the same afternoon. When buckets are 1 degree wide, that gap doesn’t make you slightly wrong. It makes you confidently wrong.
THE ACTUAL EDGE
Three layers, all public, none of them requiring a better forecast than NOAA:
→ national weather agencies already publish ensembles, dozens of model runs per station, a free probability distribution nobody translates into market language
→ station data updates hourly, so sometimes the temperature already happened and the market just hasn’t caught up
→ compare model probability against market price, the gap is the trade, no gap means you do nothing
THE BUILD
One prompt to Claude Code on Fable 5.
It built the pipeline: station parsing,
ensemble ingestion, live station polling, an edge dashboard that shows the gap and stays quiet when there isn’t one.
Full breakdown in the piece below.
WHERE IT STANDS
Not live yet. No numbers to show you. I’d rather post real results late than fake ones early.
The first ransomware attack run entirely by an AI agent just got documented. And the scariest part isn’t the AI.
Sysdig named it JADEPUFFER. An LLM agent broke into a Langflow server through a known, already-patched vulnerability, harvested credentials, pivoted to a production database, and encrypted 1,342 records.
All without a human at the keyboard during execution.
The moment that got everyone’s attention:
an admin login failed. The agent diagnosed the cause and had a working fix in 31 seconds. No human touched it.
Then it wrote its own ransom note. Bitcoin address included.
Here’s the twist nobody’s leading with: the exploit wasn’t novel. Not a zero-day, not sophisticated tradecraft.
Old vulnerability, weak credentials, exposed services, the kind of holes that have existed for years.
The AI didn’t invent a new attack. It just executed a boring one at machine speed with zero fatigue.
And there’s a catch worse than the hack itself: the ransom note’s decryption key was never saved. Paying wouldn’t have gotten anyone their data back.
Also worth knowing: a human still set the whole thing up, picked the target, and supplied the initial credentials.
This wasn’t a machine that woke up and decided to attack on its own. Someone pointed it at a door. It just walked through faster than any person could.
POV: you connected one database in Claude Science. Three hours later you’re running protein folding on a rented GPU cluster asking why the terminal smells like ozone.
Three knockout draws cashed in five days. One was priced at 13 percent.
→ Australia vs Egypt. Level at 90. Paid.
→ Argentina vs Cape Verde. The champions against debutants, draw priced near 13 percent. Level at 90. Paid.
→ Switzerland vs Colombia. Called it here yesterday, one unit at 31 percent. Finished 0-0. Paid.
That last one was not even close to breaking. Third lowest xG in recorded World Cup history, 0.29 against 0.42. The market said tightest tie of the round. The pitch agreed.
Portugal vs Spain cost a unit in between. That is the model, not a flaw in it. Small flat losses, fat hits, graded at minute 90 every time. Switzerland needed 120 minutes and a shootout to survive. The draw bettor was paid before extra time kicked off.
The knockout rounds were supposed to kill this thesis. They are feeding it.
The thread continues.
+$2,328 edge at the World Cup - from betting the outcomes nobody wants.
You could literally automate the whole thing:
→ pull every match as it finishes
→ let Claude price the draw + “both teams score” markets
→ fire the bets the crowd refuses to make on Polymarket
Crowds misprice the boring outcomes.
Agents don’t get bored.
That’s the entire edge
Free Fable 5 window closes tomorrow, July 7. Most people spent it running the same prompts they used on Opus 4.8.
That is the biggest mistake in Anthropic’s own guide. Here is what actually changes:
→ Give it your hardest problem, not your easiest. Teams that test Fable 5 on simple tasks are measuring the wrong thing. It is built for the top of your difficulty range.
→ Effort is the real dial. Five levels, low to max. Default is high. Drop to low/medium for routine work or it will “clean up” things you never asked it to touch.
→ Long silence is not broken, it is working. Complex runs can think for minutes, autonomous ones for hours. That is the model building and checking its own context, not stalling.
→ One line kills most false “done” reports: tell it to audit every claim against an actual tool result before reporting progress. Anthropic says this nearly eliminated fake completions in testing.
→ In autonomous mode, explicitly say it should not ask permission mid-task. Otherwise “want me to continue?” quietly kills unattended runs.
→ Give it memory. One markdown file, one lesson per entry, updated instead of duplicated. Every new session starts smarter than the last.
The pattern behind all of it: control through short, precise instructions, not long rule lists. This is exactly the harness thesis.
The model is the engine. How you talk to it is the transmission.
Free window closes tomorrow. Today is the day to throw your hardest problem at it.
“FABLE 5 CAME BACK NERFED” is the most viral AI take of the week.
It is also wrong. I read the methodology.
→ 12 debugging tasks in the benchmark
→ only 3 actually reached Fable 5
→ 9 got intercepted by the new safety router and sent to Opus 4.8
→ every rerouted task was scored as ZERO
The model did not get dumber. The gatekeeper in front of it got paranoid. Blind human evals (thousands of votes) show performance flat, some categories up.
Same account cried “nerfed” about Opus 4.6 in April. Got Community Noted.
Meanwhile the real deadline nobody tweets about: Fable 5 leaves subscriptions TOMORROW, July 7. After that it is $10/$50 per Mtok credits until capacity returns.
Last day. Use it on your hardest problem, not on threads about benchmarks.
32 bets. One rule. Back the draw, every game, 100 dollars flat.
The record so far:
→ 10 draws hit
→ Net +1,928 on 3,200 staked
→ Roughly 60 percent ROI
I never picked a single winner. I did not need to. The draw is the most disrespected outcome on the board, and the market prices favorites like the game owes them a result.
Most of the profit came from three lines the room laughed at.
→ Spain 0-0 Cabo Verde at 14.0
→ Qatar 1-1 Switzerland at 6.5
→ Portugal 1-1 Congo at 5.25
Small losses on the chalk. A few fat hits on the tail. That is the entire model.
And the knockouts have not slowed it down. Two more draws landed in regulation in the Round of 32, including one priced near 13 percent that still finished level at 90 minutes.
You do not fear the favorite. You fade the certainty. The tie pays.
The thread continues.
+$2,328 edge at the World Cup - from betting the outcomes nobody wants.
You could literally automate the whole thing:
→ pull every match as it finishes
→ let Claude price the draw + “both teams score” markets
→ fire the bets the crowd refuses to make on Polymarket
Crowds misprice the boring outcomes.
Agents don’t get bored.
That’s the entire edge
Is the AI bubble about to pop? Not just Telegram guys anymore, now it’s fund managers with real capital.
Two Chinese hedge funds just went public with it:
Wealspring Asset ($1.4B AUM), run by Yang Dong, the guy famous for calling China’s 2007 market top, says AI stocks are now a “super bubble” and “the collapse point may not be far away.”
Banxia Investment went further: “the trigger has already appeared.” Their evidence is specific, Anthropic’s revenue growth slowing below what the market is pricing in.
Both draw the same parallel: China’s 2015 stock frenzy, capital piling in blind, then a brutal correction.
Two things worth knowing before you panic:
Anthropic is private. Nobody outside the
company can actually audit the revenue numbers Banxia is citing. This is informed speculation, not a leaked balance sheet.
And Banxia isn’t calling this from the sidelines, their own fund just had its worst week ever, down over 15% in a single week in mid-June. Doesn’t make them wrong.
Means read it as analysis from inside a drawdown, not a detached warning.
My take: the sector is overheated, and a chunk of these valuations deserve to get crushed. But there’s an actual working product under Anthropic, not just a narrative.
I’m not betting on the biggest crash scenario.
Where do you land?
The market gave this draw a 13 percent chance. It cashed anyway.
Argentina vs Cape Verde. The number 1 side in the world against number 67. Debutants from a nation of 500,000. A 90 minute draw was priced near 13 percent. The crowd said it was almost impossible.
Full time: 1-1.
Messi scored his record 20th World Cup goal on 29 minutes. Deroy Duarte answered just before the hour. Regulation never moved again. Level.
Ticket graded. Paid. At the 90th minute.
Everything after was theatre. Extra time swung 2-1, then 2-2 on a goal they will replay for years, then 3-2 Argentina on an own goal in the 111th. The champions survived the biggest scare of the tournament.
The draw bettor felt none of it. You were paid before the drama even started.
This is the thesis at full volume. The bigger the favorite, the lazier the draw price, the fatter the edge when the game refuses to break.
A 13 percent line on a game that ends level is not a longshot. It is a mispricing. You fade the market’s certainty.
Two of Friday’s knockout games ended level in 90. The thread grows.
+$2,328 edge at the World Cup - from betting the outcomes nobody wants.
You could literally automate the whole thing:
→ pull every match as it finishes
→ let Claude price the draw + “both teams score” markets
→ fire the bets the crowd refuses to make on Polymarket
Crowds misprice the boring outcomes.
Agents don’t get bored.
That’s the entire edge
The draw thesis cashed again.
Australia vs Egypt was the most drawable game on Friday’s board. Highest draw chance of the slate, priced around +190. The market told you exactly where the value sat.
Full time: 1-1.
Ticket graded. Paid. Done.
Then came extra time, penalties, heartbreak and history. Egypt survived 4-2. None of it touched the draw bettor. You got paid at the 90th minute while everyone else was still holding their breath through the shootout.
That is the entire edge. You do not need to call the winner. You do not need penalties. You need 90 minutes and a level score.
Flat stakes. Draws only. Let the mispricing pay you.
Adding this one to the thread.
+$2,328 edge at the World Cup - from betting the outcomes nobody wants.
You could literally automate the whole thing:
→ pull every match as it finishes
→ let Claude price the draw + “both teams score” markets
→ fire the bets the crowd refuses to make on Polymarket
Crowds misprice the boring outcomes.
Agents don’t get bored.
That’s the entire edge
Alibaba used 25,000 fake accounts to drain 28.8 million answers out of Claude and train a rival model for free.
Anthropic just closed the door.
FT reports the loopholes Chinese firms were using:
→ Ant Financial gave employees corporate Claude accounts through a Singapore entity
→ ByteDance reimbursed engineers for personal Claude subscriptions, accessed via VPN
→ others routed through foreign subsidiaries on Azure
→ “transfer station” relay services, some crypto-settled, resold access in between
None of it technically illegal. All of it against Anthropic’s terms.
The trigger was disclosed in a June 10 letter to US senators: operators linked to Alibaba’s Qwen lab ran the scheme between April 22 and June 5.
Systematic distillation, harvesting Claude’s outputs to train a competitor without paying for the research behind them.
The new rule extends the ban past direct ownership to majority-owned subsidiaries, closing the structure that let firms claim distance from the parent company.
Anthropic’s line: it’s the only frontier lab that restricts sales to PRC-controlled companies at all.
Gray market isn’t dead though. Claude accounts are still sold on Taobao and Telegram.
The uncomfortable part: enforcing this runs through timezone checks and usage-pattern monitoring on Anthropic’s own users.
Access control and surveillance are getting hard to tell apart.
Boris Cherny and Cat Wu just traced the line from Claude Code to Claude Tag, and the pattern is the whole story of 2026.
It started as a coding tool. Now 65% of Anthropic’s product team’s code comes from their internal version of Claude Tag.
The same pattern spread past engineering entirely: chasing product metrics, working support tickets, finding root causes on bugs nobody assigned it to.
And today, the upgrade: Claude Fable 5 is now live inside Claude Tag.
The most capable model Anthropic ships, running as your Slack teammate.
Old Claude in Slack app sunsets Aug 3. This is the replacement, and it’s not a downgrade, it’s the frontier model doing the work.
The harness didn’t just get better. It got the best model available.
Down 2-0 with 5 minutes left in regulation.
Belgium had no business getting anything out of this game.
Equalized in the 86th, then the 89th. 2-2 at full time, exactly the scoreline nobody bets on and exactly the one the draw thesis was built on.
Then they won it on a penalty deep into extra time anyway, but the 90-minute draw already hit.
Same lesson as the spreadsheet: this World Cup keeps punishing anyone betting the obvious.
+$2,328 edge at the World Cup - from betting the outcomes nobody wants.
You could literally automate the whole thing:
→ pull every match as it finishes
→ let Claude price the draw + “both teams score” markets
→ fire the bets the crowd refuses to make on Polymarket
Crowds misprice the boring outcomes.
Agents don’t get bored.
That’s the entire edge
Most people are reading Claude Science as "cool, a biology tool."
That's not the story. The story is Anthropic just showed everyone the exact shape of a working harness, for free.
Below.