The recent breakthrough in Navier-Stokes has garnered a lot of attention to the new possibilities that emerge when agents work together.
We have been interested in this question for a while. How can a team of agents achieve more than agents working alone?
I'm excited to finally share our new work, "Self-Organizing Agent Teams Learn to Reason Together" 🧵
How to build long-horizon AI agents: behavior specs, ontologies, process supervision - my conversation with @mitch_troy, co-founder of @trybasis
01:09 Why Everyone at Basis Was Whispering to AI when @steph_palazzolo walked in
04:12 Accounting as "an Intelligence Over the Economy"
06:11 What Makes an Agent Truly Long-Horizon
08:24 Inside an Autonomous, Multi-Day Tax Return
10:19 Agents That Hand Off Like Senior Engineers
11:17 A Brief History of Agents: From ReAct to Today
12:33 Why LLMs Have No Long-Term Memory
14:13 Why AutoGPT Didn't Live Up to Its Promise
15:51 The Three Breakthroughs: Opus 3, o1, o3
17:07 Why Reasoning Models Unlocked Agents
18:23 "Let's Verify Step by Step": The Road Not Taken
20:32 Pushing Back on the METR Chart
22:09 Why Coding Agents Won First
25:14 Why Real-World Agents Are Harder
26:55 How Accountants Verify Non-Deterministic Work
29:18 You Can't Scale Tax Returns Like Math
33:16 100 Evals Pass - So What?
35:53 Right Answer, Wrong Process
36:37 Behavior Specs, Explained
39:58 How Specific Should Behaviors Be?
42:18 Context Is Runtime Training Data
44:21 Who Judges the Judge?
46:45 The Move 37 Objection
50:02 The Magic Box Mental Model
52:41 "Nothing Has Changed Since o3"
54:56 Open-Sourcing Behavior Specs with @ankrgyl@braintrust
59:45 Ontologies: A World for Agents to Live In
01:04:20 Documentation as Codebase
01:06:33 Why the Founding Fathers Were Context Engineers
01:09:05 Onboarding 300 Brilliant Alien Employees
01:11:10 Self-Improving Agent Systems
01:12:50 The Context Mistake Agent Builders Make
01:14:29 RL on Behavior Adherence
01:17:01 Will the Bitter Lesson Swallow the Harness
01:18:46 "Technical Moats Are Not Real Moats"
01:21:03 Advice for AI Builders
The future of tax is end-to-end intelligence.
Today at Basis Tax Day, we hosted leaders from 30 of the Top 100 accounting firms to explore what changes when Basis takes the first pass on complex returns.
We launched Basis End-to-End Tax for that future. So accountants get out of data entry and start with review.
The 2027 busy season will feel very different with Basis.
sample efficiency is maybe the biggest challenge for scaling agents to hours, days, etc.
we are researching problems like these every day at @trybasis, come join us!
Zooming out a bit, what we’re really doing is defining a reward function, deciding how to allocate that reward over a trajectory, and deciding how to proliferate the signal through the system.
Today, at Basis, that means using the signal to update the runtime system: context, skills, tools, the harness, model choice & specs, etc.
Couple research direction we are pursuing:
1. As agents become better at systems thinking & theory of mind over other agents, they will rapidly start closing the loop from production issue to improved deployment.
To do this well, you NEED to have sources of signal beyond outcome based evals, and we think behaviors are the missing ingredient there.
2. I suspect as agents improve their tooling to operate coherently over very long horizons, directly rewarding the model based on the process reward from the behavior might start bearing fruit.
These are the types of research directions we’re exploring at Basis. If you're interested in working at the frontier, shoot me a dm.
Based on everything I know and I've seen, @mitch_troy is pretty prescient when it comes to where AI and long-horizon agents are going. People clamor about how Basis does evals, and it's awesome to see this published out in the wild!
Out of the box, long-horizon agents struggle to accurately perform end to end work in the real economy (outside of coding) because those tasks are not easily verifiable, the data is hard to scale, and going from inputs to real outcomes can actually take many days.
Even if you had a reliable way to verify outcomes at scale (and weren’t bothered by the multi-hour iteration loops), the sheer volume of decisions by the agent that occur in a multi-hour job makes it hard to know whether performing well will generalize to production.
Over the last two years at @trybasis, we've been solving this problem by supervising the process our agents take to get to outcomes, rather than just looking at whether the outcome itself is correct.
We think this is the key to building production agents at scale.
It's what has allowed us to run agents in production that operate for hours, sometimes days, and reliably perform tasks like entire complex tax returns end to end.
Today, alongside @braintrust, we're open sourcing a standard for defining, evaluating, and eventually rewarding agent behaviors.
Thread below with all the details on how we’re scaling behaviors to close the loop for long-horizon agents.
It's awe-inspiring to pause and think about how economically complicated the modern world is and how absurd it is that it works at all. (I always think back to the @collision tweet on passion projects).
You have organizations of thousands of people, say @SpaceX, that perform millions of economic events annually and somehow that information gets compressed into legible, structured data that is TRUSTED enough by investors, banks, the IRS etc. to make billion dollar decisions.
Thats possible because of the hard-working folks who account for, then reconcile, then double check (then hire other people to audit) all of that activity.
Given the complexity about to dawn on the world as agents become economic participants, we're going to need accounting and accountants more than ever.
Excited to share our work “Multi-Agent Teams Hold Experts Back” was accepted to #ICML2026 🚀
Thank you to wonderful collaborators @james_y_zou@elb4tu@CaoHancheng Carmelo di Nolfo @sun_yanchao and Meng Cao for making my first PhD project such a fun experience!
We've raised $100m at a $1.15b valuation from @Accel, @GVteam, and existing investors to accelerate deployment of the most capable and accurate accounting agents across CAS, tax, audit, and advisory.
Basis is used by 30% of the Top 25 accounting firms and dozens more across the Top 150.
Today we're announcing the first accounting agent to complete a business tax workbook end-to-end.
Our focus on production-grade, long-horizon agents means that 12 months from now, the work Basis handles will make even this look routine.
We're looking for a few very intense people who want to build at the frontier.
US statistical agencies are probably the most impressive in the world. The fact that they schedule and perform so many data revisions, in a transparent manner, furthers my confidence that they strive to produce accurate information.