dear open labs, and i mean this with love. you keep dropping these trillion parameter giants and they're genuinely beautiful and i can't run a single one of them, none of us can, they land on huggingface to a standing ovation and then they just sit there, because the only machines that can load them are the same datacenters we were all trying to get free of.
i watched local ai turn into something real this year, people running 27b models on one gaming card, an 8gb board doing work that needed a server a year ago, whole little communities forming around whatever fits on the hardware already sitting on the desk, and that's the part that actually reaches people, the revolution nobody's putting on a keynote slide.
so here's the fair ask, give us the 40b dense, give us the 120b moe, the sizes a person with one good card or a 128gb box at home can actually hold and run and learn from and build on, because not everything has to be shaped for an enterprise with a rack to spare.
the frontier is thrilling and i'll keep cheering every open weight you ship, i mean that, but the ground is where most of us live and the ground is quietly running out of models it can lift. build for the hardware we already own. that's where the next wave comes from, it always was.
with love, from someone running yours on a laptop at midnight.
@sudoingX well, there is still a large quality difference between the models that a 24gb card can run and the ones that a 8gb card can. I see no reason why this shouldn't continue to be true in the foreseeable future. Remember that it's a three times more VRAM -- that's not nothing.
@haider1 how can anyone serious about this still think genAI lab CEOs say what they actually mean? they've yapped so much before, much of it hasn't turned out to be true. its just ceos hyping their product or industry.
small models are the most underbuilt space in local ai right now.
adoption lives at 8gb, 16gb, 24gb, and 128gb unified. that's where the people are. that's the hardware already bought and waiting.
open labs: will you meet us where we are?
@TheAhmadOsman they probably also used their own unlobotomised models to tackle the security problem.
so the fact that glm5.2 was even able to help really shows that internal models aren't that far ahead of what is released...
absolute cinema.
@ryanbrewer@bajolacurva without prior knowledge of korean, you can probably get there within 3 months if you commit to spending 8+ hours per day for that goal.