Awesome day, @satyanadella!
Windows sparked a platform shift that created a new industry for NVIDIA. Then we invented programmable shading GPUs for DirectX, which led to CUDA.
Then we partnered to bring GPU supercomputers to Azure, which helped OpenAI train GPT. That collaboration inspired us to reinvent Windows for the age of personal agents.
4 years. Thousands of engineering years between us.
So proud of what we built together.
Introducing EmbeddingGemma 2, a new open multimodal model that sets the standard for on-device efficiency.
- our first open, natively multimodal embedding model
- handles text, code, image, video, and audio tasks within a lightweight, modular 740M parameter form factor
- ideal for offline, privacy-first RAG when paired with Gemma 4
- outperforms some specialist models more than twice its size
Weights available now on Hugging Face.
Au contraire.
C'est une nouvelle ère qui s'ouvre pour les mathématiques.
Une ère où la démonstration formelle est largement automatisée et où l'accent sera reporté sur le développement de nouveaux concepts, nouvelles abstractions, nouvelles définitions, et nouvelles conjectures.
L'invention du bateau a réduit l'importance de la nage, mais a permis la découverte de nouvelles terres.
Important note regarding Grok @Bot:
Going forward, @SpaceX will use the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno and other leading APIs.
Whatever is most likely to give you the best outcome.
We're deepening our commitment to scientific research and technological discovery by committing $150 million to the Genesis Mission.
We'll also make Claude and our technical support available to more than 15 federal agencies.
https://t.co/MUmQBiC6MF
New paper advised by Yann LeCun!
"H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning"
Most world models plan in a single latent space and at one timescale, which makes long-horizon planning expensive and forces one representation to handle both low-level motion and high-level goals.
H-JEPA instead stacks JEPA world models across multiple timescales, with higher levels predicting farther into the future.
The highest level first figures out roughly where the agent should go, then each lower level turns that into increasingly concrete subgoals until the bottom level outputs actual actions.
On Visual AntMaze, this 3-level setup increases success from 18% to 73% while using less planner compute.
https://t.co/dQrd4iT1T9
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
https://t.co/7N6TPlft1P
The most important thing we can apply AI to is improving medicine and human health - very excited about this partnership with the CZI and the Virtual Biology Initiative!
Today, Isomorphic Labs joins the Virtual Biology Initiative (VBI) as a founding member.
We are helping build an open resource for the global research community to accelerate how we understand and treat disease.
Read more here: https://t.co/3FY3MS4HaR
"Decoding Looped Transformers Better for (Almost) Free"
This paper treats an early loop from a looped transformer as a rough draft and the final loop as the refined answer. It then uses the difference between them to push the final prediction further in the direction the model was already improving, with no extra training.
On Ouro-2.6B-Thinking, AIME 2024 accuracy jumps from 61.9% to 73.3%
And using only half the recurrent loops can still match or beat full-depth decoding while cutting forward FLOPs by up to 48.2%
https://t.co/ufdw5bWYpA
Reasoning from scratch, round number 6!
An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO).
00:00 Introduction
01:54 What makes a reasoning model different?
04:25 Reasoning traces and model capability
08:29 Accuracy and format rewards
11:34 Aha moments and DeepSeek-R1 training
14:41 Reasoning effort and answer length
18:38 RLHF and RLVR
23:04 GRPO vs. PPO
26:40 GRPO explained with a cooking analogy
31:43 The KL term and simplified GRPO
35:04 Loading the pretrained model
36:07 Loading the MATH training data
39:26 Sampling model responses
46:30 Computing verifiable rewards
49:55 Computing advantages
51:54 Token and sequence log probabilities
55:29 Implementing sequence log probabilities
57:37 Fixing the inference-mode error
1:02:24 Computing the GRPO loss
1:04:37 Putting the GRPO step together
1:09:19 The GRPO training loop
1:12:57 Training settings, logging, and checkpoints
1:17:24 Running training and inspecting outputs
1:19:28 Loading and evaluating checkpoints
1:22:33 MATH-500 results and training stability
1:24:05 Memory requirements and next steps
The idea of an intelligence explosion caused by recursive self improvement has been around for a long time but until very recently it did not seem imminent. Now many leading researchers think it may happen quite soon. You can read our paper about it here:
https://t.co/sgUpugjpRY
Researchers at @BroadInstitute, @UniofExeter, and beyond are already using AlphaGenome Atlas to better identify potential disease-causing DNA variants and interpret their role. 🧵
More pre-training doesn't necessarily mean better generalization?
This new paper discovers mode-hopping, where LLMs repeatedly switch between shallow pattern-matching and actual generalization, even while training loss remains stable.
For example, OLMo3-32B went from 81% accuracy to 0%, then back to 81.7% within just 40B training tokens.
On top of that, selecting an earlier 4.5T-token checkpoint instead of a 4.9T one improved GPQA transfer after math fine-tuning (36.3% vs 29.8%) and robustness to alignment attacks (53% vs 21%)
https://t.co/MZzYGCejkN
Yesterday at the White House, leaders from across our industry came together to sign the White House Accord on Super Intelligence.
The principle is simple: the companies building this technology have the primary responsibility to develop and deploy it safely, and to be accountable when they fall short.
The Accord reinforces that responsibility with robust internal controls, independent external evaluation, and board level oversight.
Super Intelligence offers an extraordinary opportunity to advance discovery, productivity, security, health, and prosperity.
The future of Super Intelligence will be shaped not only by what it can do, but by the confidence it earns. By building with ambition, rigor, and responsibility, we can help ensure this extraordinary technology expands opportunity and improves lives for generations to come.
In physics, an “impedance mismatch” occurs when two systems each work well but are poorly matched.
In this Science Blog guest post, Harvard physicist Matthew Schwartz argues that something similar is happening with AI and science. LLMs are capable at many things, but working with them as you would with a human collaborator isn’t currently the best way to elicit their scientific strengths.
To address this mismatch, Schwartz created a toolkit for exact calculations in quantitative science. Because similar calculations often emerge in very different areas of science, Claude found connections to ecology, population genetics, and a dozen other fields, and Schwartz worked with domain experts to steer it towards interesting questions.
Read more about these projects here: https://t.co/UpgSwMCz7h
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Innate is a team of 7 building capable robots for everyday life.
"The seven of us here do the work of 50," says Vignesh Anand, co-founder of @innate_bot.
We spent time with the team to capture what building means to them.
Today we’re introducing Gemini 4 Argon.
It delivers frontier performance in complex workflows across real-world software engineering, knowledge work, and cybersecurity defense with an industry-leading 1M token output limit.