You can now run Kimi K2 locally with our Dynamic 1.8-bit GGUFs!
We shrank the full 1.1TB model to just 245GB (-80% size reduction).
The 2-bit XL GGUF performs exceptionally well on coding & passes all our code tests
Guide: https://t.co/kRig50dJHR
GGUFs: https://t.co/GFO4kgxQ2q
Introducing Gemini CLI, a light and powerful open-source AI agent that brings Gemini directly into your terminal. >_
Write code, debug, and automate tasks with Gemini 2.5 Pro with industry-leading high usage limits at no cost.
You do not want to miss the MCP updates & announcements from Windows, Gradio, Caucus, Anthropic, Firecrawl, Postman, Perplexity, Cline, and more this week!
I summarized everything you need to know below.
(save for later)
We’re also launching Veo 3, our state-of-the-art video generation model.
Veo 3 lets you generate videos with sound effects, background noises and even dialogue.
#GoogleIO
A week in AI Agents is like a year in traditional software.
Here is everything that happened this week in AI Agents from OpenAI, AgentOps, LangChain, You, Flowwise, Convergence, CopilotKit, Aomni, Box, Stagehand, Cohere, & more. 🧵
(save for later)
I tried Qwen3:4b. By default, it generates responses with reasoning enclosed in <think> tags.
If you want to disable that, you’re supposed to add /no_think at the end of the prompt,
but in many cases, it didn’t seem to work properly.
Project DeepWiki
Up-to-date documentation you can talk to, for every repo in the world.
Think Deep Research for GitHub – powered by Devin.
It’s free for open-source, no sign-up!
Visit deepwiki com or just swap github → deepwiki on any repo URL:
DeepWiki is incredibly useful.
Not only does it turn a target GitHub repository into easy-to-understand documentation, but it also offers an AI chat feature.
This AI chat is smart and has been extremely helpful.
Every 20 seconds, a new Spirit Plant blooms—composed of 4 traditional Polish herbs and an ever-shifting palette of colours.
In just the first 3 days of Expo 2025, visitors have co-created thousands of unique digital plants at the Poland Pavilion. The result? An endlessly evolving garden of collective imagination.
#Expo2025 #GenerativeDesign #PolandAtExpo
🚀 Introducing NSA: A Hardware-Aligned and Natively Trainable Sparse Attention mechanism for ultra-fast long-context training & inference!
Core components of NSA:
• Dynamic hierarchical sparse strategy
• Coarse-grained token compression
• Fine-grained token selection
💡 With optimized design for modern hardware, NSA speeds up inference while reducing pre-training costs—without compromising performance. It matches or outperforms Full Attention models on general benchmarks, long-context tasks, and instruction-based reasoning.
📖 For more details, check out our paper here: https://t.co/HJiqzwnUV7
DeepSeek's first-generation reasoning models are achieving performance comparable to OpenAI's o1 across math, code, and reasoning tasks!
Give it a try! 👇
7B distilled:
ollama run deepseek-r1:7b
More distilled sizes are available. 🧵
what??
NVIDIA just dropped Project DIGITS, a $3,000 personal AI supercomputer that’s small enough to look like a Mac Mini but packs 1,000x the power of your average laptop.
Handles AI models with up to 200 BILLION parameters.
This is incredible..
M4 Mac Mini AI Cluster
Uses @exolabs with Thunderbolt 5 interconnect (80Gbps) to run LLMs distributed across 4 M4 Pro Mac Minis.
The cluster is small (iPhone for reference). It’s running Nemotron 70B at 8 tok/sec and scales to Llama 405B (benchmarks soon).
MLX Swift example can also QLoRA fine-tune Llama 3.2.
Here's the 1B fine-tuning on my iPhone 15 Pro at > 150 toks/sec. A this rate only takes a few minutes to learn some decent adapters fully on-device.