This Executive Order is an important step in strengthening America’s leadership in AI.
We look forward to collaborating with the White House to support its implementation.
https://t.co/ZwDimPrp3t
Seven new models launching at Build: let’s go!
Reasoning. Code. Image. Transcribe. Voice.
Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models
Thread 🧵
#MSBuild
Inference Optimizations Behind the MiMo-V2.5 Series API Price Reductions
Read the full technical blog: https://t.co/B5tp4tdnim
The V2.5 model family, including MiMo-V2.5 and MiMo-V2.5-Pro, is built on a Hybrid Sliding Window Attention (Hybrid SWA) architecture, which compresses KVCache storage to roughly 1/7 that of Full Attention. However, architectural advantages rarely translate directly into measurable gains in production serving. To realize these gains, we redesigned KVCache management, tiered caching, and the prefix-cache tree; addressed key challenges in SWA KVCache handling; and optimized scheduling as well as the Prefill/Decode pipeline.
Validated on real production traffic, these optimizations have increased effective KVCache capacity by nearly 5x, with server-side cache hit rates averaging 93%–95% across mainstream harness frameworks. Together with MoE configuration tuning and multimodal inference optimizations, they enable more efficient long-context inference and form part of what makes the recent API price cuts possible.
Kimi Code---kimi-datasource plugin!
It connects professional external data sources, letting you query and analyze stock prices, economic indicators, and more — all in natural language!
From data retrieval → code analysis → full report generation, now a complete closed-loop workflow.🚀
🚨 Huge if true...
Apple is reportedly paying Google to build a custom Gemini-based model that runs on its Private Cloud Compute servers to power Siri.
Apple plans to roll out the revamped Siri around March next year, possibly with an AI-powered web search feature.
🚨 BREAKING: Anthropic has developed a method to test whether AI models can understand their own thoughts.
Using a new technique called concept injection, researchers inserted specific neural patterns into Claude models, and asked if they could detect them.
Claude Opus 4.1 was able to identify injected patterns like "all caps" or "loudness" about 20% of the time, indicating early signs of AI introspection.
650 million active users who use Gemini weekly. The gap to OpenAI seems small, with currently 800 million active users. Why is that?
Google has a strength that OpenAI does not have: distribution. Google is in a strong position, and not least the Android operating system and its own Pixel phones are bringing Gemini into the hands of millions of users – once again: distribution. That is the key to success.
You have to hand it to Google – they thought very far ahead early on, and that is paying off.
Qoder 0.2.10 is here 🎉
- NES-powered Tab & Tab to Jump just got smarter.
- Run parallel Quest tasks in isolated Git worktrees.
- and more.
Here's what’s new 👇
Today we’re releasing SWE-1.5, our fast agent model.
It achieves near-SOTA coding performance while setting a new standard for speed. Now available in @windsurf.
Welcome to your Agent HQ 📍Orchestrate any agent, any time, anywhere.
Coding agents from @claudeai, @OpenAI, @cognition, @julesagent, @xai and more will become available in GitHub as part of your paid Copilot subscription.
https://t.co/3d8a2Cafqq
America’s AI Stack, powered by AMD.
Together with @ENERGY, @ORNL, @HPE and @OracleCloud, we’re proud to announce that AMD is helping advance science at scale. With AMD Instinct GPUs and EPYC CPUs at the core of the nation’s newest supercomputer and AI Factory, Discovery and Lux, the U.S. is poised to continue to lead AI innovation.