Vous sous-estimez largement le volume consommé par l’agentique.
C’est précisément pour ça qu’il y a un arbitrage prix vs task (et non pure performance). Les modèles chinois sont bons, mais souvent très lents parce qu’ils manquent de compute (d’où la hausse de DeepSeek…). Dans un monde agentique, on peut faire tourner des modèles moins chers toute la nuit. Le frontier sert surtout à orchestrer.
@Jeffmaehr@andreas_nigbur So, we're looking at ancient engineering as a secret society's blueprint? That's a wild leap from dusty ruins to modern AI architecture.
Your gaming laptop can now run the same AI models that cost companies thousands in API fees. For $0.
UC Berkeley open-sourced something called FreeToken. And the numbers don't make sense.
A $1,000 laptop with an RTX 4060 is running a 35 billion parameter model at 39 tokens per second. That's faster than what Codex runs in production.
An RTX 5090 desktop is running DeepSeek V4-Flash 284 billion parameters at 22 tokens per second. On a single GPU. At home.
3-4x faster than Ollama. 6-30x faster on prefill. One-click install. No conversions. No cloud. Your data stays on your machine.
Plug it into Claude Code, Codex, or any coding agent. It works out of the box.
A year ago running frontier models meant renting a data center. Now it means opening your laptop.
https://t.co/pYCoFnxvcx
oof.. even within anthropic, opus 5 is only selectively used for specific nerdy optimization problems
why? because it’s unusable otherwise. it doesn’t know how to talk to humans
why? because it’s trained by machines checking only whether it passed tests
@bcherny@kunchenguid Can’t wait to next version of Opus. Pulling for Claude. I used in 90% last year vs others. Fable is too expensive for me although I like the output. Dream would be Affordable Opus and Fable level output.
@notjazii Guess: Sonnet 5.5, Haiku 5. Haiku is marshmallow, Sonnet is Melon (Melon bigger than marshmallow ofc), I think they will release sonnet 5.5 and it may be 10% better at max effort, but haiku will be the new "workhorse" for claude models that probably rivals opus 5 medium
@btbytes That speed in the Gemini app sounds like a genuine productivity leap. Is the friction point really just the interface, or is it something deeper in the workflow?
The voice market did not spend this week polishing demos. It raised $280M, moved into Claude, Microsoft Azure, phone systems, partner channels, and regional infrastructure.
Here's what happened in Voice AI this week:
🚨 Hansi Flick :🗣️
« Pour être honnête, j’ai parlé avec Balde ce matin et je lui ai dit : “Je pense que tu n’es pas à 100 %, donc il vaut mieux que tu restes chez toi et que tu te reposes. Nous en reparlerons demain.” »
But there’s a catch.
A massive context window by itself means nothing.
If the AI can ingest 50k pages but screws up the one fact on page 47,382…
who cares?
The real benchmark shouldnt be:
“how many tokens fit?”
It should be:
“how accurately can it USE all that context?”
Big context ≠ useful context.