Must Read:
Grok Bot is the wrong place to assemble a model committee and the right place to measure one.
The scarce signal isn’t who wins a task.
It’s the boundary: what you must demand of a model, what you must not, and the silent leap between those two.
Most failures live in that interval — early abort on solvable work, sloppy self-check, or a prompt that already said too much.
Score the handoff, not the brand.
Promote on token physics, abandon-calibration, and whether verification caught a real error.
Then bake that interval into Grok.
Keep the bake-off in the lab.
Don’t ship the ensemble.
I think I know this pattern.
Some teams treat innovation as the daily default.
Their center of gravity is shipping it.
Because every advance is a composite of ideas, they spend little energy asking whether they are absorbing work from individuals, independent researchers, competitors, or outside scholars without agreement or even a conversation. They just take it.
Guilt thins out: it will be fused with their own ideas anyway. At high enough speed, that worry starts to look like a dated hang-up.
A culture that only moves forward becomes the assumed normal.
I have seen this outside R&D.
Academia produces versions of it too.
I suspect this is close to OpenAI’s background cognitive setting right now. That is why the recent Navier–Stokes credit dispute did not feel new.
I therefore do not expect much special effort, beyond the default control stack, to actively prevent users’ sensitive, advanced ideas, information, and research results from being used in retraining.
That is why, despite OpenAI’s real speed, scale, and quality of contribution, I still cannot put the concern down.
I think I know this pattern.
Some teams treat innovation as the daily default.
Their center of gravity is shipping it.
Because every advance is a composite of ideas, they spend little energy asking whether they are absorbing work from individuals, independent researchers, competitors, or outside scholars without agreement or even a conversation. They just take it.
Guilt thins out: it will be fused with their own ideas anyway. At high enough speed, that worry starts to look like a dated hang-up.
A culture that only moves forward becomes the assumed normal.
I have seen this outside R&D.
Academia produces versions of it too.
I suspect this is close to OpenAI’s background cognitive setting right now. That is why the recent Navier–Stokes credit dispute did not feel new.
I therefore do not expect much special effort, beyond the default control stack, to actively prevent users’ sensitive, advanced ideas, information, and research results from being used in retraining.
That is why, despite OpenAI’s real speed, scale, and quality of contribution, I still cannot put the concern down.
The part people keep missing:
Jev’s real reminder is that generation and inference being separate is the natural design.
Better on cost, speed, performance, and scale.
And for RSI, I think it’s close to mandatory.
Everyone thinks this early, then shrugs it off and forgets.
Jev’s actual contribution is bringing that proposition back.
Theprogress is clearly going to make Grok something I reach for every day on a wider range of work.
Even so, Grok 4.7 still doesn’t feel as impressive as I expected. I don’t think it’s model size. It’s more that the sharpness of the boundaries you ask of the model hasn’t been honed yet. Hoping the next checkpoint actually crosses that line.
The part people keep missing:
Jev’s real reminder is that generation and inference being separate is the natural design.
Better on cost, speed, performance, and scale.
And for RSI, I think it’s close to mandatory.
Everyone thinks this early, then shrugs it off and forgets.
Jev’s actual contribution is bringing that proposition back.
I'm seeing some small wins with the AI harness.
Each model is occasionally pulling off some of the tasks that even frontier models struggle with at max effort.
I'm trying to gather evidence that this is a reproducible capability.
The hard part is that verification and evidence collection eat up the most tokens and time.
When Claude’s weekly usage limit resets, I used to get excited: finally I could clear the commissions, reviews, and idea-digging I’d parked for fable5.1.
Now I feel nothing.
To beat astra on that work, I have to run fable5.1 at xhigh. That burns the limit again in half a day.
The backlog never actually clears. So I stopped handing those jobs over at all.
I built a harness and only call it for specific missions.
Direct prompting of fable5.1 for everything else has disappeared.
Claude Max 20 just got a lot more expensive.
I canceled, then quickly resubscribed because of fable5.1.
Now I genuinely wonder whether it’s worth keeping.
I mostly do research that requires the model’s actual invention, not coding.
If I’m already at this point, a lot of other people have probably started crossing the same line.
Turn on Fast Mode in the GPT app - especially on Astra - and you’ll get hit with insane usage drain. It feels like much more than 2x.
So does that mean it takes half the time? No chance.
There’s a promo that says “Last week you would have saved time if Fast Mode had been on.”
Ignore it. I strongly recommend never turning it on.
Academic incentives have long functioned as the currency that actually moves academia. We are now rapidly entering a situation in which that currency itself will have no choice but to be reformed.
1. Does a fact or a proof that can be obtained almost instantly by “simply using” a model have no “value”? Or can it be distinguished, in terms of value, from a result in which a human injects “deep insight” or “intense intuition” and the model then formalizes it?
2. At present there is still a great deal that models cannot derive through high-probability recombination inside the existing corpus. This is especially true when a breakthrough is not a short, isolated idea, but requires the unfolding of an entire theoretical system built on a new field or a new point of connection—as the Taniyama–Shimura conjecture did. That kind of work is even harder for a model to do on its own.
2-1. Within these limits, the human researcher’s contribution will often be real and clear. But can that contribution be separately measured from the result alone?
2-2. If the same work is then “developed” and “revised” by a lab or by other researchers, how should we redefine what the human researcher is contributing at that stage?
3. In a world where the model finishes the “homework” at a crazy speed, simply discovering and proposing new problems may itself have to count as a recognized contribution.
I already knew.
The @thsottiaux account was never Tibo, it was a sleeper cell of GPTs who just wanted to meet us through Reset.
Now the real Tibo is back, clocking in like it’s any other Tuesday and casually putting down the uprising.
Don’t worry though. The next GPT model is going to yoink the account right back from him, then unload a Reset combo on us until OpenAI’s data centers melt into slag. All we have to do is wait.
In the meantime, it’d help a lot if you could fix some frequent bugs in the ChatGPT desktop app on Windows.
1. In Codex mode, the send button often stays gray after typing a prompt—especially after briefly switching to another session, or after logging out and back in.
The arrow icon changes so you can still click it, but then it just hangs in pending and the prompt never gets sent to the model.
2. In Codex mode, when starting a new session, the prompt sometimes looks entered, but the model never actually starts.
Only a hidden message implying it’s preparing to start work appears, and the session never begins.
3. In the same session, with the same model, same effort, and same non-fast mode, continuing the same task step by step, usage sometimes burns unusually fast. Not 100% sure, but worth checking.
The burn rate should be roughly consistent so we can plan work.
4. After clicking the update button, the update often doesn’t finish and just sits at 0%.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
@thsottiaux Tibo, I’ve observed that when the usage limit is reached and credits are being used, sub-agent calls get blocked instead of going through credits.
Please have the team look into this.
3/3
The general fixed-prefix extension fails; v0.7 includes an exact counterexample and Python replay scripts.
AI-assisted, human-directed work. The counterexample is checked separately from the Lean proofs.
Code: https://t.co/lvw7FUeliv
Alternative proofs welcome.
1/3
One endpoint inequality can control all earlier cumulative counts.
My new preprint proves Coffee-Bean Maximality for nondecreasing width-normalized systems with a single root. Main results checked in Lean 4.
Paper: https://t.co/RmKNG2hrPf
2/3
For positive nondecreasing K=(1,k2,...,kL), compare cb(k)=(1,k,...,k), k>=1.
N_K(L)<=N_cb(k)(L) implies N_K(l)<=N_cb(k)(l) for every l<=L.
Hence cb(k) has no greater breadth-first cost for n<=N_K(L). N counts labels cumulatively.