Building LLM 2.0 & what comes after LLMs
Founder @ TEQO Labs
Founder @ TEQO circular composites (€15M+ industrial buildout)
Founder of SERVEO, IaaS sold 2022
I can genuinely say I'm building what I believe will be called LLM 2.0, although that definition hasn't landed yet.
What's interesting is that I already see a roadmap beyond it.
The next few years won't be about bigger models.
K3 doing inference at 0.32 tok/s proves why efficiency matters. Bigger isn’t always better. It’s like parking an entire library in your living room just to read one book. We need smarter models, not just bigger ones.
@victormustar I'm not sure if you are right. Accordingly to some papers distillation loses up to 45,9%. REAP / prune you effectively lose experts. Lets see what the community will bring us in sense of quality on this model.
If it becomes a real product with actual users, absolutely. But if people submit essentially the same prompt, we shouldn't be burning the same amount of compute every single time. Aside what do we learn from it if you'd just replay that. We should optimise for efficiency so we can solve more real-world problems with the same resources.
Impressive demo, but what's the point of burning 1.3M tokens for a showcase? The future of AI isn't about spending more compute—it's about doing more with less. Efficiency should become the metric that matters, not who can afford the biggest inference bill.
I applied Matt's prompt, ran it for a little over ten hours. About 1.3 million tokens on Opus 5. It created the game. I also made a video of it using Opus 5. Here's the video (and in my reply to this is a link to the actual game so you can play).
@0xBaltar@TechMDAI Personally I choose Enterprise gear yet for your desktop usage that MikroTik is more then sufficient. In some DCs you can find MikroTik in 50% of the racks. Yet the ports tend to overhead. If you scale use that SB7800 or SB7790 (soft needs to run your server with second one)
@__tinygrad__@mkratsios47 I think Mistral being behind the frontier proves this isn’t business as usual. The question isn’t whether distillation exists, it’s what actually happened. Chinese labs closed the gap far faster than almost anyone expected.
@antirez A couple of posts ago you said we shouldn't talk down others. I actually agree. Let's focus on what we can build here and be grateful for what Europe already does well, like healthcare.
On AI, I'd rather build a better alternative from 🇪🇺, which I'm working on. Happy to talk
@WilliamBryk Impressive. I count 32× Lenovo ThinkSystem SR680a V3s in that cold aisle. Is the rest of the suite filled as well, or is the cluster spread across multiple locations?