@levelsio Give it more specific context? Clearly if you ask without context you get a normie answer. With a unique context you get a normative answer for that context, so should be more unique.
Just to clarify this benchmark. This is an apple to oranges comparison.
- Cerebras is fast for batch size 1 but slow for batch size n.
- GPUs are slow for batch size 1 but fast for batch size n.
I get >800 tok/s on 8x H100 for a 405B model for batch size=n.
Cerebras' system is great for batch size=1 inference, but usually, this is only effective for personal use or latency-critical applications since it has too high a token/$ cost for deployments.
Some people point to agents, but agents usually also need to do parallel calls over multiple data points. With my open-source-derived system for agentic tasks, I can get >100 ktok/s across 128 GPUs for Qwen 32B.
I want to point this out because people in the comments seem confused or skeptical. The numbers are real, but the comparison is not transparent about the batch size differences.
This is common in many other marketing benchmarks, so when you see big gaps like this it is usually not a fair comparison.
From SQL to AI: Modern LLM prompting is declarative programming reborn. 'Give me a CSV with these columns...' is the new SELECT statement. Define your output structure, let the model work out the implementation.
1. Keep answers short, sweet, and concise.
2. Provide insightful and "lowkey goated" responses.
3. Make answers first principled:
- Break concepts down to fundamentals.
- Build up from these fundamentals for every query.
4. Follow all previous instructions while implementing the first-principles approach.
https://t.co/ElmitvPdmd for @Neovim is awesome.
I hated the UI and workflow in Cursor - reaching for a mouse feels like failure.
Thanks @yetone and @aarnphm!
Talked to my friend @paul_bridger who works in ML/AI, he was spot on, saying this:
“AI will not replace people, people with AI will replace people without.”
OMEGA Labs is holding our first ‘Garage AGI’ series on Spaces where we dive into the humanity of AGI development. We’re betting big on Bittensor and @parshantdeep is going to discuss why with @_sshahid_ and @paul_bridger
https://t.co/fRw2n089Wq