An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
We believe it will be a major step for scientific reasoning. https://t.co/iP6cyheZ7i
Google DeepMind just introduced Gemini Robotics 2.
It's a single VLA that unlocks physical dexterity across different end effectors, hands or grippers, from one model checkpoint.
Apptronik's Apollo 2 humanoid, with the 22-DoF SharpaWave hand, ties knots and seals a ziplock bag
Today, we are releasing Inkling-Small.
Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.
https://t.co/BtYNcpkDRA
Fine-tune it on Tinker today, or chat with it in text, image, and audio on Tinker Playground.
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta!
🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇
🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!
Check out the configuration details in our official API docs: https://t.co/smCwQZMeiq
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.
Announcing Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities.
https://t.co/XUiX1CouKz
We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here...
You can now use it to scan repositories, track findings across runs, verify fixes, and add security checks to CI/CD.
This is an early release, and we're listening to your feedback as we continue improving it.
Interesting.
Grok 4.6 releases around August 7. This will be the 1.5T model with significantly improved SFT & RL.
Grok 4.7 will be the 2.1T model released a few weeks later. This will be better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency.
Today, we are announcing a series of updates that give customers frontier-grade security at half the cost.
MAI-Cyber-1-Flash is our first cybersecurity model, built ground up to find the most challenging vulnerabilities in complex code bases. When combined with MDASH, it delivers world-class performance at 50 percent of the cost of leading models.
We are bringing this capability to market through Project Perception, a complete agentic security offering grounded in real-world signals and security workflows. Teams of specialized agents work together to simulate attacks, detect and triage/investigate, and fix and remediate.
This is the benefit of building the harness, context/signals, and action space separate from one model family. By combining specialized models and data with the right agents, tools, security context, and harness, we can advance the frontier of cost to outcome.
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
Today, we’re releasing Ling-3.0-flash—a hybrid-reasoning MoE model built for production-scale agents.
124B parameters. Just 5.1B active per token.
With 1/8 of the total and 1/12 of the active parameters, it matches or beats our 1T flagship model on most benchmarks shown.
Announcing Fugu-Ultra v1.1 🐡
We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedback, and trusted Fugu with real work.
Today, we’re releasing Fugu-Ultra v1.1 → https://t.co/hhO6qTawgb
Upgraded to incorporate the latest frontier models, resulting in stronger performance across every benchmark shown, including gains of up to 7.9 points over v1.0, with particularly strong results on ProgramBench and Terminal Bench 2.1.
Fugu-Ultra v1.1 is more capable across coding, agentic tasks, and advanced reasoning, and available at the same price as Fugu-Ultra v1.0
The frontier keeps moving, and Fugu keeps getting better.
We’re developing Project Camellia, a long-term AI infrastructure project in Effingham County, Georgia.
There is significant work ahead. Here is the approach guiding that work:
- Georgia families will not subsidize this project. OpenAI will pay all infrastructure and electric-service costs; Georgia PSC rules prohibit passing them to existing ratepayers. It will also be one of the first data centers designed to proactively reduce demand during grid stress.
- The project is being designed with closed-loop cooling that recirculates water, similar to a car radiator. Once fully built, ongoing water use is expected to be comparable to a similar size office building.
- The project is expected to create thousands of construction and permanent on-site jobs, generate hundreds of millions of dollars in state and local tax revenue, and prioritize local contractors and businesses.
- OpenAI will provide $80 million in community benefits over the life of the project, shaped by local priorities, plus up to $71 million in Codex credits for eligible Georgia college, community college, and technical school students.
- We’ve been meeting with state and local officials, community leaders, schools, and regional partners. A public open house this week and continued conversations will help shape how we benefit the community. An annual independent public audit will help hold the project accountable.
https://t.co/oxxP92V4aP
Today we are Introducing BTL-3.
A 27B open-weight agent model built for agentic coding, structural tool use . The complete thing fits in one 8.39GB file under 2.5 bits per parameter smaller than an 8B model in fp16, and retains 92.2% of the 27B itelligence
BTL-3 is trained for the loop real agents live in: reason, act, inspect the result, recover, continue. It handles single, sequential, and parallel tool calls and knows when the right move is no tool call at all.
HumanEval: 95.12% pass@1
BFCL v4 AST: 88.5% (full 1,240-case set)
Multiple tool calls: 95.5%
Tool-call abstention: 91.2%
262K context architecture
Two editions, both open today.
BTL-3 is the maximum-quality checkpoint, for Transformers and vLLM.
BTL-3 Compact is the entire model in one standalone 8.39GB GGUF. No base download. No reconstruction. One file, one command, a running agent.
Compressing 27B this far normally destroys a model. Standard quantization couldn't do it, so we built the stack ourselves: packed AVQ2 decoder tensors, affine INT4, measured precision islands, packed vocabulary matrices, rank-32 output correction, behavioral repair. 2,416 tensors byte-verified at export.
Then we tested whether the agent survived. On a fresh sealed 100-turn tool-contract gate, Compact retained 92.2% of teacher-correct behavior 100% on single, parallel, sequential, and abstention calls.
43 tok/s generation on an RTX PRO 6000. Fully local. Nothing leaves your machine.
BTL-3: https://t.co/ddZWWr6i3o
Compact: https://t.co/6URHEBGJgG
Runtime + source: https://t.co/MjXQR6koKt
Apache-2.0 model. MIT runtime.
We are excited to announce that AMD and @AnthropicAI are expanding our strategic partnership to accelerate the development and deployment of next-gen AI infrastructure. Tune in at 9:30am PT tomorrow to hear from Dr. @LisaSu from the #AdvancingAI keynote stage!
✅ Up to 2 GW of AMD Instinct MI450 Series GPUs in AMD Helios
✅ AMD has committed to make a strategic equity investment of up to $5B in Anthropic
✅ Deep engineering collaboration across Claude, ROCm and AMD Instinct
More on the news: https://t.co/zA0vQHTT1c
Today we announce a new paradigm for quantum control. By integrating reinforcement learning with quantum error correction, we enabled a quantum computer to continuously adapt to drift, stabilizing the system during computation. This improved logical stability 3.5x. Learn more: https://t.co/U1jTSmFrdQ