we got the full Kimi K3, 2.8T params, running on 80x RTX 5090s.
20 tok/s single stream, day one, untuned. Last week we took GLM-5.2 from 30 to 110 tok/s on this same fleet. This number will climb.
A first for open weights: frontier intelligence served with zero HBM, the scarcest silicon in AI. Just GDDR7 gaming cards, plain ethernet, and the official MXFP4 weights, nothing requantized.
The most powerful open model on Earth, on the most abundant GPUs on Earth. Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it.
@Kimi_Moonshot
Thinking Orbs, animated component library
6 types, optimized for large and small sizes, light & dark mode, playground, zero dependencies
https://t.co/PdGxwi6Kws
npm install thinking-orbs
Collaborated on this with @a_brinza who did a lot of work
Fun interactive science app ideas | Part 10
Built an app for exploring different types of butterflies
It's surprising how calm and peaceful the gentle wing movements feel
Everything is generated with code. No 3D models
Code
Kimi K3
🚀 Introducing Qwen3-Omni — the first natively end-to-end omni-modal AI unifying text, image, audio & video in one model — no modality trade-offs!
🏆 SOTA on 22/36 audio & AV benchmarks
🌍 119L text / 19L speech in / 10L speech out
⚡ 211ms latency | 🎧 30-min audio understanding
🎨 Fully customizable via system prompts
🔗 Built-in tool calling
🎤 Open-source Captioner model (low-hallucination!)
🌟 What’s Open-Sourced?
We’ve open-sourced Qwen3-Omni-30B-A3B-Instruct, Qwen3-Omni-30B-A3B-Thinking, and Qwen3-Omni-30B-A3B-Captioner, to empower developers to explore a variety of applications from instruction-following to creative tasks.
Try it now 👇
💬 Qwen Chat: https://t.co/rJ5DUlBJBr
💻 GitHub: https://t.co/VHHp1jWcHT
🤗 HF Models: https://t.co/WsP8vyWltp
🤖 MS Models:
https://t.co/sGnbZRfxVz
🎬 Demo: https://t.co/v2WrZvCKD6
The first direct observation of gravitational waves captured the merging of two black holes which emitted 36 septillion yottawatts of power (3.6×10⁴⁹ watts), greater than the combined power of all light radiated by all the stars in the observable universe.
[🎞️ SXS simulation]
ChatGPT's new Image Generation dropped less than 24 hours ago
Here are 15 great examples of what you can do now, some limitations—and a hidden trick to get instant access if you're still waiting!
1. Life-like photos:
CtrLoRA can adapt a base ControlNet for image generation with just 1,000 data pairs in under one hour of training on a single GPU!
It reduces learnable parameters by 90%, making it much easier to create new guidance conditions.
https://t.co/oBnh7WR2rH
We updated Phi-3 mini in our June release, with the enhancement in instruction following, reasoning in MMLU (70.9)/GPQA(30.6), and better long context. Share with us your feedback on the new models
https://t.co/Yg4Awrvu1K
https://t.co/4UIpTHnZ9P
Here’s a neat paper by Barnett et al. (@DeakinA2I2) that outlines 7 failure points in building a RAG pipeline over your data.
🚫 Missing content (did not index it)
🚫 Missing in top-k retrieved set
🚫 Missing in reranked set
🚫 Not extracted (in context but LLM couldn’t use)
🚫 Wrong format (e.g. JSON)
🚫 Incorrect specificity (not at the right level of granularity)
🚫 Incomplete - the synthesized answer only answers part of the question
We've posted a lot about this on the @llama_index side but this diagram nicely covers a lot of the aspects. If I were to add a few, I’d add failure points during the query understanding/rewriting phase (particularly if you’re building agents).
Check out the paper: https://t.co/PPH9PLq3gj