We’re releasing EmbeddingGemma 2, our first natively multimodal open model engineered for on-device embeddings.
Built on the Gemma 4 architecture and released under an Apache 2.0 license, it goes beyond text to unify images, video, audio, and code in a single embedding space.
Local Laya moggs Jev at @grok 4.7-built Tetris 🧩
An open-weights System One model called Laya, beat cloud-based Jev at playing Tetris by making decisions 11 times faster, running locally on a 16GB MacBook Air!
Run AI models locally -> https://t.co/RbcCOIgVkj
Meet Husky: a Model-Specific Inference (MSI) engine up to 4.5× faster than Apple's MLX
Woof, Underdog's Pareto frontier model, now runs up to 730 tokens/sec on a MacBook
Finally local models are as fast & capable. Try it now in https://t.co/hAWKvlClUC - your personal private AI
This looks... interesting 👀
Just discovered Hemmingway-1, a model based on Qwen3.8-27B that promises to sound and write like a person than any other model around.
I am going to test it out!
https://t.co/BWIYM3gclI
MiniCPM5-2B beats Qwen3.5-4B in 34% less time 👾
We gave three models the same real-world tasks and compared their attempts using Q4_K_M quants
Tests:
• Build a playlist within strict timing rules
• Trace a checkout failure through server logs
• Check stock and draft a replacement with tool calls
Outputs:
• MiniCPM5-2B: 3/3 tasks, 23.0s total
• Qwen3.5-4B: 3/3 tasks, 35.0s total
• Gemma 4 E2B: 2/3 tasks, 29.5s total
MiniCPM and Qwen cleared all three, but MiniCPM finished faster, meanwhile Gemma found the failure’s cause but broke the JSON-only rule
Operating System powered by Qwen 3.8 27B at 1950 tokens/sec!
here is what 1,950 tokens/second @Alibaba_Qwen's 3.8 27b actually looks like on @cerebras:
i wrote a minimal python web server that turns cerebras inference into a live operating system.
zero apps on disk.
when you double click an icon:
1) python proxies a raw SSE stream from qwen 27b at 1,950 tok/s
2) calculator compiles & mounts in 11s
3) full canvas paint studio with brush engine compiles in 10s.
at 2,000 tokens/second, software is just an on demand hallucination that runs instantly.
the model weights ARE the operating system runtime.
what else would you build at 1,950 tokens/second?
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
1/6
WTF... Someone literally built a massive collection of open-source, in-browser tools that require no sign-up to use. 🤯
It includes tools across multiple categories:
→ Design & Graphics
→ Development
→ Productivity
→ Privacy & Security
→ AI
→ Education
...and many more.
No accounts. No sign-ups. Just open your browser and start using them :)
Follow me for more amazing AI, Coding & Web Dev insights 💎
Big news for AI on a budget. GLM 5.2 Colibri int4 is a Mixture of Experts model that runs entirely on your CPU. No GPU needed. It's fast, efficient, and opens up advanced language capabilities to anyone with a standard computer. This is a game changer for offline AI.
You run Kernels, not models
The model is just a graph
The Inference Engine serves as a scheduler, optimizer, and executor
But the actual work? That happens in the Kernels
- MatMul Kernels
- Attention Kernels
- RMSNorm Kernels
- KV cache Kernels
- Quantized linear Kernels
- Sampling Kernels
- Fused “please don’t write this back to memory 9 times” Kernels
Same model, same GPU, same VRAM
Wildly different performance
Because one stack is using optimized fused Kernels that understand your hardware
And the other stack is playing hot potato with tensors through 47 tiny launches and pretending the GPU is the problem
Bad Kernels make people say:
“this model is slow”
Good Kernels make people say:
“wait how is this running locally?”
This is why Inference Engines and the Kernels implemented within them matter
The model is the recipe
The hardware is the kitchen
The Kernels are the knives, pans, burners, and the chef not cutting onions with a spoon
Most people benchmark models
The real ones benchmark the Kernels underneath