Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
Sakana AI Teams With NVIDIA to Advance Open Model Innovation from Japan
We're announcing the next phase of our collaboration with NVIDIA. We're bringing NVIDIA's open model stack, including the Nemotron family, into Sakana Fugu, our multi-agent orchestration system.
https://t.co/bsxjRkux0a
Rather than relying solely on scaling individual monolithic models, our approach focuses on collective intelligence. Sakana Fugu operates as an intelligent orchestrator behind a single API, dynamically selecting, coordinating, and combining the strengths of multiple models for each task. This architecture keeps our system modular, adaptable, and resilient.
As a natural next step to expand Fugu's capabilities, we're integrating NVIDIA Nemotron as a specialized agent, complementing the frontier and open models Fugu already orchestrates. Nemotron helps demonstrate how open models become far more useful when orchestrated within agentic systems rather than used in isolation.
This collaboration creates a reinforcing cycle. Fugu gains a deeper pool of specialized capabilities, while NVIDIA can evaluate how its models perform when coordinated within complex, multi-step workflows. These real-world signals can continuously improve both the models and the orchestration layer.
By combining Sakana AI's Japan-born collective-intelligence approach with NVIDIA's open models and accelerated computing, we aim to shape a future of AI that is modular, collaborative, and open by design.
Meet kbd-1.0-codex-micro, built with @work_louder.
Map the buttons and joystick to your workflow, and keep your pinned chats in view.
Get yours before stock returns 410.
DeepSeek-V4 can now run locally with Unsloth GGUFs! 🐳
Run lossless DeepSeek-V4-Flash on 168GB RAM.
3-bit works on 110GB Mac, RAM, VRAM setups.
We improved the chat template. Run via Unsloth Studio or llama.cpp.
Guide: https://t.co/mCZZkpa95X
GGUF: https://t.co/wT4e1JMjp8
How do we make LLMs faster and lighter? Don’t force the GPU to adapt to sparsity. Reshape the sparsity to fit the GPU! ⚡️
Excited to share our new #ICML2026 paper in collaboration with @NVIDIA: "Sparser, Faster, Lighter Transformer Language Models". This work introduces new open-source GPU kernels and data formats for faster inference and training of sparse transformer language models:
Paper: https://t.co/3Avj8N8iYO
Blog: https://t.co/SqFkkKvkbd
Code: https://t.co/PHSzMq8pg0
While LLMs are undoubtedly powerful, they are increasingly expensive to train and deploy, with a large part of this cost coming from their feedforward layers. Yet, an interesting phenomenon occurs inside these layers: For any given token, only a small fraction of the hidden activations actually matter. The rest approximate zero, wasting computation. With ReLU and very mild L1 regularization, this sparsity can exceed 95% with little to no impact on downstream performance.
So, can we leverage this sparsity to make LLMs faster? The challenge is hardware. Modern GPUs are optimized for dense matrix multiplications. Traditional sparse formats introduce irregular memory access and overheads that cancel out their theoretical savings for GEMM operations.
Our contribution is twofold:
1/ We introduce TwELL (Tile-wise ELLPACK), a new sparse packing format designed to integrate directly in the same optimized tiled matmul kernels without disrupting execution.
2/ We develop custom CUDA kernels that fuse multiple sparse matmuls to maximize throughput and compress TwELL to a hybrid representation that minimizes activation sizes.
We used our kernels to train and benchmark sparse LLMs at billion-parameter scales, demonstrating >20% speedups and even higher savings in peak memory and energy.
This work will be presented at #ICML2026. Please check out our blog and technical paper for a deep dive!
Most motion papers tailor one controller to one specific task. This year at SIGGRAPH, our research team asks: can motor control itself be pretrained and reused?
Generative Pretrained Controllers, or GPC, turn motor skills into a vocabulary of discrete tokens and train a transformer-based generative controller through next-token prediction. Just like GPT, the same pretrained controller can then be fine-tuned to solve new tasks.
Trained on 600+ hours of motion, GPC runs in real-time inside a physics simulation, producing natural and physically grounded behaviors for interactive control.
We took a 30B model and split it in two to write tokens in parallel instead of one at a time.
Introducing Nemotron-Labs-TwoTower: a diffusion language model from NVIDIA Research adapted from Nemotron-3-Nano-30B-A3B. Here’s how it works: one half holds the context, the other writes the tokens, with both reusing the pretrained model instead of training a new one from scratch.
We found it kept 98.7% of the original model’s quality at 2.42× faster generation.
Introducing Claude Sonnet 5, our most agentic Sonnet yet.
It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models.
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work.
https://t.co/OoM83SyISN
NVIDIA Metropolis Blueprint for video search and summarization (VSS) 3 is here.
Now your coding agent can analyze massive live streams and libraries of videos with a simple natural language prompt. Here's what's new:
- 16 new agent skills: Search, summarize, alert, report, review clips. All from natural language prompts.
- One unified open source repo: Source code, Docker and Helm deployment profiles for fast, easy deployment.
- Multi-video reports and Nemotron 3 Nano Omni: Insights across video and audio at scale.
- 3D multi-camera tracking: Production ready + #1 SOTA for smarter scene understanding.
Try VSS skills 👉 https://t.co/XvKJ0Kb8VV
NVIDIA announces NVIDIA Halos for Robotics, the industry’s only full-stack, open robotics safety system.
As robots move alongside people and equipment, Halos gives machines that sense, decide and act in the real world a single common safety architecture. https://t.co/PACFHrlrT2
Fugu stands shoulder-to-shoulder with leading models like Fable and Mythos across the industry's most rigorous engineering, scientific, and reasoning benchmarks.
Read the full blog: https://t.co/2ZJbdWqCUj
Beyond Bigger Models: Why are Orchestration Models the Next Frontier
Progress in AI has been driven largely by giant, monolithic models. But the most powerful systems of the future will be collaborative ecosystems.
Today, this orchestration is no longer just a technical optimization. It has become a geopolitical and operational imperative.
For an organization or a nation, relying on a single company's model for critical infrastructure, finance, or governance is a material vulnerability. This risk is no longer a hypothetical possibility, but a reality.
As we have seen with recent export controls imposed on models like Fable and Mythos, access can disappear overnight.
Collective intelligence is the practical hedge against this concentration of power. Because Fugu orchestrates an underlying pool of swappable agents, it simply routes around vendor restrictions.
By orchestrating the world’s models, we are delivering the resilient blueprint required for true AI sovereignty.
GLM-5.2 can now be run locally!🔥
The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size).
Run on a 256GB Mac or RAM/VRAM setups.
GLM-5.2 is the strongest open model to date.
Guide: https://t.co/bI7FeeKHDd
GGUF: https://t.co/BMkxswdj5N