Congrats to everyone with papers accepted to #NeurIPS.
I am thrilled to share that we successfully withdrew one of our #ICLR submissions, because we caught a serious evaluation error before submitting the #slop.
Story: We used Codex/Claude Code to help design and debug LLM evaluations on modest GPUs. When we hit OOM errors, Claude Code suggested several fixes, including reducing the max output-token budget. Tired of auditing every suggestion, we approved it.
Two months later, while analyzing the results, we discovered that this “small” change substantially altered the evaluation outcomes. The reduced generation budget became a major confound and made our method look much better than it actually was.
We thought we had made significant progress.
BUT it was, to a large extent, an evaluation artifact.
AI can accelerate research. It can also accelerate confounding. And thanks to AI who helped us confirmed the error...
NOTE: Image is AI-generated for illustration only; all numbers are made up but the story is true.
Can your AI agent tell when another agent is lying?
Our MindGames report is accepted at NeurIPS! 🎉
944 agent submissions. 76 teams. ~30K games.
Built on TextArena, with game data + top competition agents to test against. Paper and code link below ↓
My summer internship at Apodex has come to an end, and I’m excited to see Apodex 1.1: Scaling Agentic Intelligence for Complex Work released today. It was a pleasure to contribute to the work and work with the team.
A personal research perspective I’m taking away from this summer (shaped by both the experience and many conversations with friends/labmates) is that, as base models become increasingly capable, the bottleneck may be shifting from “Can the model do this?” toward “Can the system provide the right information, level of verification, and computation at the right time?” Below is my own view.
From this perspective, I find multi-agent systems especially interesting as a form of context scaling. Different agents can explore separate parts of a problem in parallel, surface complementary evidence, challenge intermediate conclusions, and compress their findings into a better decision context.
The direction I’m most excited about next is connecting this idea with self-evolution.
In particular, I’m interested in environments that natively support multi-agent learning and capability acquisition. Rather than the generator–solver setup, where one agent creates an environment and another solves it, I’m interested in environments where multiple agents can interact, collaborate or compete, generate new experiences for one another, and gradually acquire reusable capabilities.
Going forward, I’ll continue exploring this direction through my PhD research while also staying on as an intern at Apodex. I’m excited to approach these questions from both the academic and applied sides.
Congrats to everyone involved in the release!
How to make our academia better?
❌ Destroy it by submitting 100 papers to NeurIPS
✅ Construct it by submitting your thoughts to AI-Native Academia workshop @ NeurIPS
https://t.co/PZD8amQZja
Deadline: Aug 29, 2026
Thank all the organizers
@denghui_zhang@Jianing9810@ManlingLi_@VITAGroupUT@gramchurn 🎉
We now release the StateM runbook to achieve 88.8% on Terminal-Bench-2.1 at https://t.co/sDKL5an2yc! Everyone can try!
Note: it was adapted from the 95.3% GPT runbook, so...
StateM: harness scaling for reliable agents
A runtime that gives agents durable state, checked transitions, and recoverable runbooks. On Terminal-Bench 2.1, it lifts GPT-5.5 to 92.1% and DeepSeek-V4-Flash to 88.1% for under $15.
This week is turning out to be all about the harness breakthrough.
Haven't verified this myself yet, but with harness scaling, GPT-5.6-SOL broke through the previous ceiling on Terminal Bench 2.1 and hit 95.3%.
Matches GPT5.6-Sol-Max with Deepseek-V4-Flash (88.8%).
They say applying this to open-source models gets them to frontier-level performance. Definitely need to test this out.
Archive ⬇️
Is model scaling the only source of agent improvement?
We @henryqin1997@YaxinLu1997@VITAGroupUT@VictorKaiWang1 are glad to share our work: We reach 95.3% Raw Accuracy, or a $15 Frontier Run (88.8% with DeepSeek-v4-Flash, matching GPT-5.6 Sol Max), on Terminal-Bench 2.1 via Harness Scaling.
Human effort -> Non scalable ❌
Agent direction? Local optimization ❌
Human direction + agent action = better workflow control✅
StateM:
1. Open-source model + StateM -> Close-source frontier, everyone can try!
2. Training free, significantly improve frontier agents.
3. Zero-transfer-cost on same family of agents. [1/8]🧵
📄https://t.co/GjM1VOzWUN
🌐 https://t.co/KapEB3kjXr
🚀 New paper: Attention is All You Have.
56,804 agent skills exist. Your model only has ~100 reliable trigger slots.
Every install dumps a skill’s description into the permanent system prompt. They all fight for the same scarce attention. The long tail dies. Your team’s private playbooks compete with random public ones. Token tax every message. Dilution. Reliability collapse.
Installation was never the right primitive. It bundled three things that don’t belong together: content, persistence, and auto-triggering. Only the last one needs to live in the prompt.
The new skill protocol separates them.
- A path addresses any skill, subtree, or collection
- Reading it is using it, zero residency
- :save vendors a clean copy into your project’s git tree so you own and adapt it
- :install adds one .gitignore-style line. That’s the only thing that costs prompt space.
Any agent that can read files + run commands becomes a client with a single instruction file. Optional free hub for search + ranking across the entire corpus.
Install less. Use more.
Attention is all you have.
A collaboration between @adalagent and @VITAGroupUT (Dr. Atlas Wang)
🚨A novel way to do RL in LLM post-training!
Inspired by our previous path-not-taken work (https://t.co/hKe4qkgiXL), we dig deep into the learning trajectory of RL and find that optimizing singular vectors (i.e., rotation) of weight matrices suffices for good performance in RL. The resulting “isospectral optimization” reaches matched scores with substantially fewer training steps.
Great work from @zhu_hanqing666 and the co-authors!
People keep asking me: what's different about optimization in RL?
Seemingly nothing — the pre-training stack just works (Adam, even SGD 👀 @saagnikkk).
Bringing some answers from my last work (sorry for the delay — been cooking 🚀).
We introduce ISO: Isospectral Optimization: an RLVR-native optimization stack.
Built on one simple observation, spectral inheritance: RLVR can reuse the base model's spectrum and acquire new behavior purely through the singular frames.
🧩 Offline: ISO-Merger — consolidates RL experts into one model with no data, no rollouts, no OPD. Checkpoints only.
⚙️ Online: ISO-Optimizer — a drop-in wrapper on AdamW / Muon that matches AdamW's accuracy with ~2.7× fewer steps on Qwen3-8B-Base.
📄 https://t.co/ty1FrAveA5
🌐 https://t.co/xzXf6V8ukO
🧵👇
🚨Academic publishing is becoming AI-native before its governance is ready.
📣Announcing the NeurIPS 2026 Workshop on AI-Native Academia: Authorship, Peer Review, and Conference Governance under AI.
📍 Atlanta, Dec 12/13 (in-person)
We are honored to host an all-PC-chair panel on "Redesigning AI Venues Under AI" — with the program chairs of NeurIPS, ICLR, ICML, CVPR, MICCAI & MIDL: Yisong Yue (@yisongyue), Kyunghyun Cho (@kchonyc), Sharon Li (@SharonYixuanLi), Mert Sabuncu (@mertrory), and Humphrey Shi (@humphrey_shi), moderated by Atlas Wang.
💡 We are also thrilled to welcome our invited speakers: James Zou (@james_y_zou), Kyunghyun Cho (@kchonyc), Hima Lakkaraju (@hima_lakkaraju), Hiromu Yakura (@hiromu1996), Yian Yin (@yian_yin), Lin Peng (@LinPengFin), and Nihar B. Shah.
We invite work across the whole pipeline: AI-assisted authorship, prompt-injection attacks on AI reviewers, LLM-review detection, human-AI co-hallucination, citation integrity, corpus contamination, venue policy.
4 pg / 9 pg tracks, non-archival. Under-review papers welcome.
📅 Deadline: Aug 29 (AoE) | Notification: Sep 29
📝 Submit: https://t.co/hW3nbi3lFa
🌐 https://t.co/zuUAM9mokG
🙌 Organized by Denghui Zhang (@denghui_zhang), Jianing Zhu (@Jianing9810), Gopal Ramchurn (@gramchurn), Manling Li (@ManlingLi_), and Atlas Wang.
Submissions are warmly welcome — we look forward to seeing you in Atlanta!