NVIDIA LPU supports 3 types of disaggregated inferencing:
1. Rubin Prefill + LPU Decode for the fastest interactivity
2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve
3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve
For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.
This is catastrophic for Anthropic. They continuously deceive their users in every possible way.
Misleading communication, usage limits, anti open source behavior, aggressive stance towards people who use Claude outside of Claude Code...
At some point it becomes a clear pattern.
@varunram That's the trap here. I would bet that more than 20% of all token traffic is from OpenAI.
"User traffic" includes free tier with Composer models and casual use.
They are pissed.
@dr_l_alexandre Et pendant ce temps, certains parlent de retraite à 60 ans, de taxer les super-profits, d'interdire les data centers...
La classe politique française vit en l'an 1995. Terrible.
Lack of concrete experience is the most common failure mode of stereotypically smart people. The stereotypical "smart kid who doesn't know what the hell they're talking about." As a kid it gets you a head-shake with a smile. As an adult there is no smile.
Comme disait De Gaulle, les états n'ont pas d'amis, que des intérêts.
Oui, je suis un pragmatique et j'ai le sens du business. Les valeurs morales sont utiles pour rallier les foules.
L'émotion n'a pas sa place dans les négociations. Seule la froideur des intérêts prévaut. C'est ainsi que la France doit dominer.
I have integrated Gemma 4 31B from @cerebras in Codex.
Now, I can use a model at 1800 tokens / second to review my architecture and act as a teacher. While 5.6 Sol Max stays the orchestrator, Gemma is very useful when it comes to actually explain to you tradeoffs, quirks and more.
Soon, they'll have Qwen 3.8 27B. I can't wait.
No Astra this week sadly. But we will have some fire Neo lab updates later this year.
Hark will unveil new hardware this fall
So will OpenAI.
Astra still on track for early September. But I feel OpenAI has waited too long just in time for Anthropic to reclaim the throne shortly after.
It's the 2-year anniversary of the best blog post of 2024, "You Are NOT Dumb, You Just Lack the Prerequisites" by @lelouchdaily.
https://t.co/bEjR1K7Cij
@Proton_Pass What's your opinion on generating a 4 to 7 words password with diceware ?
Like, for example :
Sadness Liver Paradigm Binoculars Pragmatic Ribbon