Long reasoning without the quadratic tax: The Markovian Thinker makes LLMs reason in chunks with a bounded state → linear compute, constant memory and it keeps scaling beyond the training limit.
1/6
Introducing linear scaling of reasoning:
𝐓𝐡𝐞 𝐌𝐚𝐫𝐤𝐨𝐯𝐢𝐚𝐧 𝐓𝐡𝐢𝐧𝐤𝐞𝐫
Reformulate RL so thinking scales 𝐎(𝐧) 𝐜𝐨𝐦𝐩𝐮𝐭𝐞, not O(n^2), with O(1) 𝐦𝐞𝐦𝐨𝐫𝐲, architecture-agnostic.
Train R1-1.5B into a markovian thinker with 96K thought budget, ~2X accuracy 🧵
On-policy self-distillation has strong pass@1 but shows "collapse" in pass@k. The reason: the same model A. generates the rollouts, B. provides the demonstrations, and C. acts as the teacher. 3 layers of self-confirmation that compound into collapse. Check out Andrei's post!
Hi!
Just landed in South Korea 🇰🇷for ICML and joined the Search Technology Meetup! If you’re around, let’s grab a coffee!
I also have two ICML workshop papers:
📄 OCR: https://t.co/wYwe3tQ679
📄 Sycophancy: https://t.co/DuRu8REXv1
1/ 🧵 Meet Tapered Language Models (TLMs):
Modern language models (transformer, recurrent, memory-based) are a stack of *identical* layers. A uniform parameter distribution across depth, inherited from the 2017 transformer and rarely questioned.
Turns out it's leaving free performance on the table.
Perplexity drops 16.28 → 14.44 (same params, same compute) 👇
Excited to share our latest work: Loss Smoothing for Stable Adaptation Under Distribution Shift.
TL;DR: When adapting a model, don’t abruptly switch to the new objective. Gradually smooth the transition from the old objective to the new one.
Preprint: https://t.co/gkJd29pqYe 🧵
1/n
I'll be presenting The Markovian Thinker at this year's CRL Symposium! Excited to share our latest work and learn about all the exciting research from our lab over the past year. Hope to see many of you there.
We are super excited to announce @ChandarLab 7th Annual Research Symposium! 🔥
Join us on July 23–24 for two days of talks on deep learning, NLP, reinforcement learning, continual learning, and AI for science, plus a keynote by @mengyer from NYU!
Event is hybrid!
New paper from the lab:
Ask a base LLM to pick a random weekday, and it might output "Wednesday" 80% of the time. We show that this can be fixed and that probabilistic calibration is actually a trainable capability.
@ZyphraAI@AMD Very exciting to see Markovian Thinking used in ZAYA1-8B. Scaling test-time compute to millions of tokens within 32K context, with strong gains from a <1B active parameter model, is exactly what we hoped this idea would enable. Congrats to the Zyphra team!
https://t.co/ajVK8SMujV
Excited to see that Markovian Thinker contributed to Zyphra's strong release 🚀. Their Markovian RSA: markovian thinking (carrying forward bounded-length reasoning tails) + RSA (recursive self-aggregation) boosted test-time compute to be on-par with larger reasoning models. 1/
Excited to see that Markovian Thinker contributed to Zyphra's strong release 🚀. Their Markovian RSA: markovian thinking (carrying forward bounded-length reasoning tails) + RSA (recursive self-aggregation) boosted test-time compute to be on-par with larger reasoning models. 1/
Diffusion world models can help test and improve robot policies before running them on real robots.
But can the choice of latent space make the WM more faithful?
We show that semantic spaces beat reconstruction spaces on task relevant metrics.
https://t.co/BwHZk7ciUQ
🧬 New paper
Scientific datasets evolve as science evolves. With proteins, new sequences get added, annotations get corrected, and noisy entries get curated out.
Introducing CoPeP, a continual-pretraining benchmark for protein LMs.
Details 🧵
1/n
Streaming Reinforcement Learning (RL) is a huge challenge: transitions are used once and discarded immediately. This makes agents extremely sample-inefficient. But what if we could "squeeze" more information out of every single frame?
Check out our latest paper!
‘The Markovian Thinker’, developed by our lab, has been accepted at @iclr_conf! This work achieved long reasoning without the quadratic attention tax LLMs reason in chunks with a bounded state, achieving linear compute, constant memory and scaling beyond its training limit!
New work from our lab, accepted @iclr_conf : "The Expressive Limits of Diagonal SSMs for State-Tracking"
We give a complete characterization of what diagonal SSMs can and cannot compute on state-tracking tasks and the answer is deeply connected to group theory.
🧵👇
Congrats to Prashant (@prashantg_17), Davide (@DavideBald42296 ), Quentin (@qfournier2), and Sarath (@apsarathchandar) on CADmium, a new method that rethinks text-to-CAD to generate high-fidelity 3D models! Read their blog post: https://t.co/c3b6U3bIWl
Can LLMs become CAD designers?
Check out “CADmium: Fine-Tuning Code Language Models for Text-Driven Sequential CAD Design”, which is now published in Transactions on Machine Learning Research (TMLR), and led by @prashantg_17, @DavideBald42296, and @qfournier2!
Alongside @NeurIPSConf in San Diego, the satellite conference NeurIPS Mexico City is taking place, with several Mila student-researchers taking part. Two of them presented their research today.
SaharDastani (@sonia_dt98), PhD student at ETS/Mila, presented “TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses” and Saba Ahmadi (@Saba_A96), affiliated researcher at UdeM/Mila, presented “The Promise of RL for Autoregressive Image Editing.” Congratulations!
With all the ICLR 2026 drama, we’re sharing some insights on the review and rebuttal process from ICLR 2025 & 2024. You might find them useful for your own rebuttal!
https://t.co/aticQ1n1iB
The data of scores before and after rebuttal is also available: https://t.co/cRv6nsNZ2O
🚀 Announcing GroundCUA, a high-quality dataset for grounding computer-use agents. With over 3M expert annotations spanning 87 desktop apps, we use our new dataset to train state-of-the-art grounding models, namely GroundNext-3B and GroundNext-7B.
👇 Thread
We show a phase transition for optimal data curation: For strong models, concentrating on difficult samples drives further improvement (LIMO). In contrast, weaker models benefit from the conventional "More is More" where broad data exposure is essential to learn core capabilities
After nearly 3 years since our NeurIPS paper, SOTA architectures are now adopting NoPE. Kimi Linear uses NoPE for all full-attention layers (not a RoPE hybrid).