Nvidia just took Claude Opus 5 from 30% to 100% on ARC-AGI-3 without touching the model.
The only thing that changed was the harness — memory management, a supervisor loop, and Nvidia's Agentic Variation Operators framework. 100.00 RHAE across all 183 public-set levels, in 6,624 environment actions versus VISTA's 7,542.
A 70-point swing from scaffolding.
Now read that against every model-vs-model leaderboard you've stared at this year. If the wrapper is worth 70 points, those charts are measuring plumbing, not intelligence — and "which model is smartest" becomes close to unanswerable in any agentic setting.
Credit to Nvidia for putting the caveat in their own post: this "should not be interpreted as a controlled ablation."
For anyone shipping agents, the takeaway is uncomfortable and useful: your edge was never the weights. It's the loop you build around them.
#AI #AIAgents #MachineLearning #NVIDIA
The biggest crypto week in three years had almost nothing to do with crypto.
Bitcoin ran ~24% to $79,400 — its best week since March 2023. Here's what actually drove it:
The US 30-year hit 5.33% on Aug 18, the highest since 2007. Fiscal supply, 3.4% CPI, and AI-related corporate issuance all crowding the long end at once.
Treasury blinked. Wednesday it doubled per-issue buybacks from $2B to $4B. Thursday Bessent went further on CNBC: "it could be more than $4 billion per issue."
The relief lasted about a day. Tens went from 4.64% straight back above 4.70%.
Now look at what the BTC move was made of: $1.24B liquidated in 24 hours, $1.06B of it shorts. Eighty-five percent short-side.
That's not accumulation. That's a squeeze, lit by a fiscal-credibility scare.
Stocks still closed the week red. Gold finished Friday at $4,680.
If your thesis this week was "adoption," you were long the right asset for the wrong reason.
#Bitcoin #Macro #Bonds #Trading
Foreign holdings of US Treasuries: the "dump" narrative vs the actual TIC data.
End-June 2026: total foreign holdings $9.299T, down $72.1B MoM but still +2.3% YoY. Recent peak: $9.489T in February.
The real story isn't liquidation.
It's composition.
Official holdings (central banks): $3.778T — range-bound between $3.6–4.1T for years.
Stagnant.
Private holdings: growing.
UK at $939.9B (+9.8% YoY) is mostly custody for global funds. Belgium (+$10.5B) = Euroclear. Cayman ($453B) = hedge fund domicile.
China: $633.4B, lowest since 2008, –13.4% YoY.
But this is a decade-long diversification from a $1.3T peak, not a sudden exit. Japan's –$26.4B drop is FX intervention mechanics, not policy.
Net TIC flows in June: +$133.5B inflow across all asset classes.
The structural shift: price-insensitive official buyers → price-sensitive private buyers, while US debt issuance outpaces both.
Foreign share of total Treasuries is ~30–33%, down from >50% at the GFC.
That's not a crisis.
It's a repricing of who sets the marginal bid — and at what yield.
RUNECLAW v11-8B eval is in. Closest yet, and the clearest signal that 8B has hit its ceiling.
v11 vs v8/v10 on the same eval_prompts_v2 yardstick:
R:R traps caught: 3/3 (incl. the 1.18 boundary)
Textbook trades approved: 3/3
Verdict accuracy: 73.5 (v8 still leads at 82)
Headline gates: 2/5
First time trap discipline and textbook approvals coexisted — v8 and v10 could only hold one at a time. Daily-loss and loss-streak gates recovered too.
But cooldown, stale-data and macro-event gates stayed broken for the third straight generation. And the most telling line in the whole file:
"Meme Coin Guard: 4.6% > 4.0% — PASS"
Math correct, verdict backwards. That's not a data-volume problem. Three recipes, three winning profiles, none holds everything — the model is trading behaviors off because it can't hold them all.
Deployment: no change. v8 stays on scan/thesis, v10 on chat. v11 goes to the registry and deploys nowhere.
Next: v12 on the 14B base. Two data fixes going in — time-gate weights 2→3, and the meme guard becomes contrastive (45% legitimate passes at 1.8–3.9% ATR, rule spelled out in both branches). 3,000-sample validation: zero auditor flags.
Bar for v12: gates 5/5, traps 3/3, textbooks 3/3, zero mislabeled comparisons. ~4–5 days on the 5090.
#RUNECLAW #AlgoTrading #CryptoTrading #LLM #FineTuning #Qwen #MachineLearning #BuildInPublic #HumanoidTraders