Meta-harnesses have been on a tear recently, it's pretty tempting to direct our quant research efforts into this area. But I do wonder how this plays out as intelligence scales - the customization required for harnesses to work well will continue to go down, which begs the question is it worthwhile for us to devote resources into this, vs spending time on our core trading models. Perhaps I'm overly cautious from our experience working on evolutionary algorithms, where we built a custom variant of AlphaEvolve, researching improvements to map elites and incorporating multi armed bandit theory, which gave us useful insights but ultimately I think these "10,000 agent swarms" will supersede much of that work.
We sent a swarm of AI agents to solve Karpathyโs NanoChat benchmark, and they crushed SoTA in 3 days.
We built our own harness on top of a graph database to do auto-autoresearch, each iteration learning from mistakes to do research better.
The team wrote >15,000 entries. ๐งต
Someone asked me why creating a simulation/environment for trading isn't straightforward. One I can speak to from my experience at HFTs like Virtu - in some asset classes like options, many firms never actually built a simulator for a few reasons:
1. For options, there are many more inputs needed to construct the "state of the world" beyond delta one, for example all the classic black scholes inputs (each of which has its own sub-system to produce), adjustments (a typical approach for trading a portfolio of options is to adjust each option's theo value based on your existing positions for risk management), quoter (liquidity adding system), taker (liquidity removing system), etc. And you have to account for how all these interact with each other, e.g. simulating what the taker saw conditioned on what the fitter gave you conditioned on what the adjustment was at that time. Many firms did not build for this when they started, and tacking this on later became too difficult.
2. Path dependency - the adjustments point above touches on this - for a portfolio of options your trading is highly path dependent. Every position you hold impacts your adjustment on every other option in the entire vol surface (across strikes and across expiries), and even across vol surfaces in similar underliers (e.g. SPY ETF vs ES futures). This means every trade you make highly impacts what your next trade will be.
3. Then there's the classic challenges with building a simulator in any asset class - things like market impact / how other participants react to your actions - market makers love to "dime"/"penny" each other, some trading lingo referring to if market maker A improves a bid from 100.00 to 100.01 โ market maker B will improve to 100.02 โ A will improve to 100.03 โ etc. Or what happens if you place a giant 10K share bid when the market size is usually just 100 shares per side - this has ties to the commonly talked about "reward hacking" in RL - if your environment doesn't account for how a 10K bid will make the market go berserk, an LLM tasked with improving the PnL of a liquidity adding strategy will certainly exploit this.
Lilian Weng's post (https://t.co/HqOTAj9CdG) provides some good insights into this. One challenge with self improving meta harnesses (essentially the "outer loop that questions the recipe") that especially resonates is her point about preventing diversity collapse being critical for open-ended research, since the best path may initially look worse under the current evaluator.
This is a common problem in similar loops for quant research - how do you prevent the population from collapsing into variants of the same solution, which can happen at both the autoresearch level and the auto-meta-research level. For the most part right now you still need human imposed frameworks, which don't work particularly well in my domain, but I think a trend we'll see is smarter models โ simpler harnesses and greater ability to question effectively in the outer loop.
To get to ASI we likely need auto-meta-research, not just auto-research.
Auto-research hill climbs within the current recipe. Minimize pretraining loss, maximize post-training evals.
Auto-meta-research defines new objectives. An outer loop that searches across paradigms. Outside deep learning, maybe even outside gradient descent. Not just scaling transformers + RL.
The inner loop optimizes the recipe. The outer loop questions the recipe.
I've become convinced that creating an environment for agent swarms to operate freely in is the bet to make for quant trading, rather than spending more time designing how they should search. Agent swarms with latest models can already solve Navier Stokes - I think as intelligence scales, we'll see increasingly less human imposed frameworks (like evolutionary algorithms) and more of a "let them rip" approach.
At our quant firm we did quite a bit of research earlier this year on evolutionary frameworks for trading. There have been many publications over the last year inspired by Deepmind's AlphaEvolve/FunSearch to write code for trading strategies, for example CogAlpha (https://t.co/moFMLHhCcJ). A lot of the research goes into things like map-elites, how do you create the right balance between exploration vs exploitation, how do you optimally choose which paths to exploit further. Given we're a two man team, we have to make the right bet with our time/resources, and I think continuing research in this area isn't the right move.