Want to teach Gemma to master chess?
Check out this awesome community project showing how to fine-tune Gemma 4 12B on your own data, 100% locally!
Running text, images, and audio on just 8GB VRAM makes custom models more accessible than ever.
I accidentally discovered that Gemma-4-26B-A4B is way better at writing human sounding content than every other model out there - including frontier models like GPT 5.5 and Sonnet 4.6.
I'm not sure why this is - it's kind of crazy how slopified these big expensive models are and for some reason, Google's open source model sounds a lot more natural and follows writing instructions better.
WTF?
It's very cool that Apple shipped a 20B parameter on-device.
You can't put 20B parameters in RAM at any reasonable precision. To make it work they are using pretty exotic architecture by today's standards.
A small model predicts from the query (or prompt) which experts to load from Nand into RAM. The key distinction from a typical MoE is that you do this once per query and then generate all the tokens with the same experts (instead of switching the experts for every token).
Gemma 4 MTP just got officially merged into llama.cpp
This means you can use Gemma 4 QAT + MTP for a lightweight + super fast setup. Excited to see what the community builds with it
https://t.co/1te7tgdi2H
Google releases Gemma 4 QAT. ✨
You can now run Gemma 4 at 3x less memory with near original performance.
Quantization-Aware Training (QAT) makes it possible to run Gemma 4 26B-A4B on 16GB RAM.
GGUFs: https://t.co/wQgEocxUId
QAT Guide: https://t.co/Nsm1yeGEHx
Opus 4.8 is a total disaster. The problem is not the model per-se, they have Mythos and can anyway train a better model. The problem is: what is it happening inside Anthropic, at the management level? Since this is a product failure. If there was a technological issue NOT delivering is better.
Meet Gemma 4 12B!
A unified, encoder-free multimodal model designed to bring high-performance intelligence directly to your laptop, and released under an Apache 2.0 license.
Bridging the gap between edge efficiency and advanced reasoning. Here is what’s new with Gemma 4 12B: 👇
llama.cpp now has an official website: https://t.co/vztdUpdBWL
Our goal is to make local AI accessible to everyone, and improving the user experience is a big part of that. On the new landing page you’ll find a single-line cross-platform installer. The installation provides a single unified `llama` entrypoint which you can use to run/serve models and interface with 3rd-party agentic applications.
While oriented towards simplified user experience, the new `llama` application also provides all the advanced functionality of the existing llama.cpp tooling with which experienced users are already familiar. Also note that all GGUF models that you might have already downloaded with llama.cpp in the past will be automatically available to use without downloading again (they are stored in the common HF cache on your machine).
We have many improvements in the pipeline both at the UX and at the engine level and we plan to iteratively ship new things over the coming months. One of the main focuses will be seamless integration with local-friendly 3rd-party agents (such as Pi). In the meantime, we’ll continue to listen for feedback from the community and adjust accordingly, so keep letting us know what you think and need.
Einstein on (not) using NL for invention: "The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought"
Would you like to join the research effort on JEPA and World Models easily?
After a full year of hard work, we’re excited to finally release stable-worldmodel:
an open-source, scalable platform built to accelerate JEPA & World Model research!
📄: https://t.co/gnxGvens5A
Good morning, Mneme AI version 1.3 just released with new control center, lock screen and action button actions (on iOS 18) and a new widget! Please let me know what features you want to see next!