After all the discourse surrounding the restraint of AI’s progress, this stands as the final statement.
AI changes the limits. Institutions and character decide whether those new limits become freedom or a prettier cage!
Long context is priced in memory, not tokens.
KV cache per token is 2 x layers x KV heads x head dim x bytes.
A 70B with GQA (80 layers, 8 KV heads, 128 dim, fp16) holds about 320 KB per token. One 128k request parks about 40 GB in HBM.
That's why your batch size collapses the moment users paste whole repos.
GQA, FP8 KV and paged attention are the real long context features.
i just automated our release and QA process!
using Grok Bot i made two Team Bots: sandcastle, my release manager, and poteto, my engineer bot. i added them to Slack and made a new channel for managing releases.
first i tell sandcastle that i want to do a cut. it'll DM all the contributors with links to their PRs so they know what goes in the next release and can DM the bot back if they have any blockers or objections.
sandcastle then kicks off the build, watches it automatically, and starts up an automated fuzz swarm. usually this is 10+ agents running on Grok 4.7 xhigh to run the build, click around and use it like a real user (using our verification skills and feature map), some are directed while a few are "chaos monkey" clicking things randomly to see what breaks.
whenever it finds an issue, i tell it to @ mention my engineer bot poteto, to spin up a Cursor Project to own triaging and fixing all the high pri issues. depending on the severity i may ask it to cherry pick the fix into our release branch and cut a patch release, or if it's a pre-existing issue i still fix it but include in the next release instead
that's one more tedious task that Grok Bot now handles for our team
An agent talking to a business is the easy part. HTTP already does that.
The hard part is the mandate. Which actions can my agent take without asking me, up to what dollar amount, and how does the merchant verify that before it ships?
Without signed, scoped, revocable delegation, every checkout turns into a chargeback argument about what the agent was allowed to do.
And the business gets to talk back. Every merchant reply is untrusted input landing in my agent's context.
The spec that wins solves authority, not transport. @SierraPlatform@Meta
Today we’re announcing Personal Agent Protocol — an open standard @Meta and @SierraPlatform are developing along with industry partners at @Genesys, @instinct, @RocketOTD, @Shopify, @stripe, and @Walmart. It will help define how personal agents interact with businesses and is open for anyone to implement.
Read more: https://t.co/9cWCVlCv8H
If you tune your prompt against the same 200 eval cases for a month, those cases stopped being an eval.
They became training data. You just did the gradient descent by hand.
The score keeps climbing. Production doesn't move.
Hold out a slice you never look at. Rotate in fresh traces from real traffic every week.
Read failures on the dev set. Report numbers on the held out one.
One embedding space for text, images, video, audio and code, running on the device.
That deletes a whole layer of glue. No captioning video just so you can search it as text. No separate index per modality.
The catch is the modality gap. Contrastive models tend to park each modality in its own cone, so a text query can rank a mediocre caption above the clip you actually wanted.
Measure cross modal recall on your own data before you merge the indexes. @Google
We’re releasing EmbeddingGemma 2, our first natively multimodal open model engineered for on-device embeddings.
Built on the Gemma 4 architecture and released under an Apache 2.0 license, it goes beyond text to unify images, video, audio, and code in a single embedding space.
Foveated rendering, now for diffusion.
Graphics stopped shading every pixel equally a long time ago. Generative models kept paying full price for blank walls.
Attention is quadratic in token count, so coarse tokens on flat regions save more than linearly.
The real question is the layout. When the detail prior misses, small text in the background is the first thing to melt. @GordonWetzstein
Diffusion models spend the same compute on a blank wall as on a face. But you often know in advance where the detail will be.
Introducing Level-of-Token (LoT) Diffusion: we turn that knowledge into a multiresolution token layout, with fine tokens where detail is needed and coarse tokens elsewhere.
1/9🧵
Research taste doubling every 3 months is the most important curve in that forecast, so read the baseline.
Most of their experts haven't worked at a frontier lab. Beating them is real. It isn't beating the people choosing frontier runs.
And taste only compresses the search. A model that picks the right experiment still waits on the training queue to find out it was right.
Judgment shrinks the tree. Compute sets the clock. @pzeroresearch
How fast is AI's research taste improving?
We find that the experimental research taste of frontier models has doubled every ~3 months since December 2025. The best model, Opus 5.5, now exceeds our expert human baseline. Our human experts are experienced researchers, but most haven’t worked at a frontier lab.
Why measure research taste? In the AI Futures Model, it largely determines how quickly artificial superintelligence is reached once coding is fully automated.
49B active sets your cost per token.
1T total sets your hardware bill.
At FP8 that's about a terabyte of weights before any KV cache. An 8x H100 node doesn't hold it. You need B200 class memory or two nodes.
Open weights at this size means open to whoever already owns the cluster.
Still a good thing. Just know who it's for. @MistralAI
Meet Mistral Large 4, aka Le Chonk.
• 1T parameters, natively multimodal. 49B active.
It is the best open weights model from US or Europe on aggregated benchmarks.
• State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding.
• Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure.
• Available to all via API today. Working with cybersecurity partners privately.
Open weights release end of October.
49B active sets your cost per token.
1T total sets your hardware bill.
At FP8 that's about a terabyte of weights before any KV cache. An 8x H100 node doesn't hold it. You need B200 class memory or two nodes.
Open weights at this size means open to whoever already owns the cluster.
Still a good thing. Just know who it's for. @MistralAI
Meet Mistral Large 4, aka Le Chonk.
• 1T parameters, natively multimodal. 49B active.
It is the best open weights model from US or Europe on aggregated benchmarks.
• State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding.
• Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure.
• Available to all via API today. Working with cybersecurity partners privately.
Open weights release end of October.
Speculative decoding speedup is a property of your traffic, not your model.
The draft only pays off when the big model agrees with it. JSON, boilerplate code and copied context get accepted in long runs. Open ended prose gets rejected token by token.
Same stack, huge win on one workload, barely moves on another.
Benchmark it on your own logs, not the paper's.
Temperature 0 doesn't make your model deterministic.
The server batches your request with strangers. Batch size changes the reduction order inside the matmuls, and floating point addition isn't associative.
Same prompt, different neighbors, different tokens.
If your eval reruns once and calls it a regression, you might be measuring traffic.
Reconstructing a scene from video as code is a harder world model test than predicting the next frame.
A video model can produce plausible pixels without knowing anything about mass, contact or occlusion.
Code has to commit. Get friction wrong and the ball rolls off the table, and you can diff the render against the real clip.
Inverse graphics turns "understands physics" into something you can actually grade.
Do coding agents understand the dynamics of the world well enough to reconstruct it from video?
We introduce 4DCodeBench to evaluate this ability through 4D inverse graphics.
https://t.co/QlkzTPZpbi
<🧵>
Your tool's error messages are part of the prompt.
"500 Internal Server Error" teaches the agent nothing, so it retries the same call three times.
"limit must be 100 or less, got 500" fixes it in one step.
Most agent retry loops I see aren't model failures. They're bad error strings.
Write errors for the model, not the log.