How can we train small agentic models that are highly capable of terminal use and coding?
Announcing OpenThoughts-Agent + OpenThinkerAgent-32B, the strongest Qwen-3 based open-data agentic model: 44.8% avg across 7 agentic benchmarks! (1/n)
ByteDance has just released Seed 2.1, which achieved outstanding performance on two of our benchmarks: FrontierCS (https://t.co/1Oo1J0GttV) for frontier computer science problems, and WorldBench (https://t.co/BRNM4c5tsR) for multimodal knowledge.
Model Card: https://t.co/0H0wtCTw97
Today, agents execute isolated tasks. Tomorrow, agents will steer complex decisions across long horizons.
Introducing CEO-Bench, a first step to measure "Steering Intelligence." In CEO-Bench, agents are asked to run a simulated startup for 500 days.
https://t.co/9gNWFALTKK
Can a small academic team build a strong text-to-image model using only public datasets?
Introducing i1: a simple, fully open recipe for strong text-to-image models
Today’s vision benchmarks suggest VLMs are nearing saturation, but real-world visual understanding is far from solved.
Introducing WorldBench: 2,000 hand-written, human-verified VQA questions focused on visual diversity and designed to be challenging for frontier models. Gemini-3.1-Pro leads with just 64.0% accuracy. (1/10)
UEval: 1,000 questions, 10,417 rubric criteria, 8 tasks.
A challenging benchmark for the next generation of unified models.
Paper: https://t.co/Jk7xe4xPNQ
Dataset: https://t.co/aldylO9Uyl
Code: https://t.co/5p7BxEWZKS
Work led by Bo Li, and with @DavidYin0609, @wenhaocha1, @XingyuFu2.
Stronger Normalization-Free Transformers – new paper.
We introduce Derf (Dynamic erf), a simple point-wise layer that lets norm-free Transformers not only work, but actually outperform their normalized counterparts.
Last night, @agupta and I hosted a great dinner with 14 professors at #NeurIPS2025 from leading academic labs across the US, and many cited compute in academia as "abhorrent". Out of curiosity I just pulled these stats. This is insane. To do meaningful AI research today you need at least 1 GPU/student. Likely 8+ to be honest. The best university (Princeton) is at 0.8 GPUs/student. Stanford is at 0.14 GPUs/student. Marlowe (Stanford's "super cluster") has only 248 H100s for the whole CS Dept to use. Every frontier lab has >100k.
This needs to be fixed.
Excited to share our lab’s first open-source release: LLM-Distillation-JAX
supports practical knowledge distillation configurations (distillation strength, temperature, top-k/top-p), built on MaxText
designed for reproducible JAX/Flax training on both TPUs and GPUs
Excited to share our new work: “Learning to See Before Seeing”! 🧠➡️👀 We investigate an interesting phenomeno: how do LLMs, trained only on text, learn about the visual world?
Project page: https://t.co/9mQt3qnckL
Still remember earlier this year, I tried very hard to figure out why vLLM can give very different outputs than huggingface models even with greedy sampling (not sure if it is fixed now). Twisting batch size or number of gpus also makes the output different
Today Thinking Machines Lab is launching our research blog, Connectionism. Our first blog post is “Defeating Nondeterminism in LLM Inference”
We believe that science is better when shared. Connectionism will cover topics as varied as our research is: from kernel numerics to prompt engineering. Here we share what we are working on and connect with the research community frequently and openly.
The name Connectionism is a throwback to an earlier era of AI; it was the name of the subfield in the 1980s that studied neural networks and their similarity to biological brains.
https://t.co/lrJioBmpbT
How do we navigate a growing collection of post-trained LLMs?
In Delta Activations: A Representation for Finetuned LLMs, we propose a compact embedding that encodes the post-training signal.
Try the interactive model navigator 👉 https://t.co/EhxzGihVu5
Check out our new paper “Generative Modeling of Weights: Generalization or Memorization?” — we find that current diffusion-based neural network weight generators often memorize training checkpoints rather than learning a truly generalizable weight distribution!
Can diffusion models appear to be learning, when they’re actually just memorizing the training data?
We show and investigate this phenomenon in the context of neural network weight generation, in our recent paper “Generative Modeling of Weights: Generalization or Memorization?"
While we're at "recognizing evals", here is a legendary vision paper from CVPR 2011
It shows that it's quite easy to classify which dataset an image comes from (39% acc, vs random=8%).
The point being, every dataset having its distinct signature should be a default assumption.
What an exciting #ICRA2025 week!✨
Catch our Oral & Poster this Wednesday in the Robotic Foundation Models Session — with @DavidYin0609, @zekaiw04 and @Dantong_Niu
Oral: Wed, 3:20 — 3:25pm, Room 314
Poster: Wed, 3:50 — 4:25pm, Room 314
Project: https://t.co/bGLv4oUJ7u
Please come to our oral and poster presentations on Wednesday — I’ll be there along with @zekaiw04 and @Dantong_Niu to present our work!
Oral: Wednesday, 3:20 — 3:25pm, Room 314
Poster: Wednesday, 3:50 — 4:25pm, Room 314
Paper: https://t.co/VO2ZVwtFj4
Have you wondered whether LLMs could directly execute robot tasks?🚨
Excited to share our work, *RoboPrompt*, a framework that enables off-the-shelf text-only LLMs to directly predict robot actions with ICL demonstrations without training! @berkeley_ai
https://t.co/bGLv4oUJ7u
I had the same experience with ChatGPT. Once it misunderstands a question or gets something wrong, it will never go back to normal with more clarification.
Pro tip: open a new tab each time
LLMs Get Lost in Multi-turn Conversation
The cat is out of the bag.
Pay attention, devs.
This is one of the most common issues when building with LLMs today.
Glad there is now paper to share insights.
Here are my notes:
🧵1/
Everyone says toxic data = bad models.
But what if more toxic data could help us build less toxic models?
Our new paper explores this paradox. Here’s what we found 👇