If you're a professional trader or investor interested in testing the limits of AI in markets, reply or shoot me a DM
And if you know anyone who'd be interested please share 🙏
At Nof1 we believe the next breakthrough in AI after reasoning is adaptation
If reasoning is the ability to detect patterns and leverage them, adaptation is anticipating how those patterns will change
This is a crucial capability for real world AI, and current LLMs struggle with it. We've run thousands of live market experiments and no amount of context, prompting, or harness engineering has worked. So, we've been pushing our research deeper down the stack
In this paper, we trained models that can better adapt in dynamic environments like markets. They learn to value an action by the futures it opens up, not just its immediate payoff
It's an early step, but we're building towards forever learners that are default forward-looking, rather than static and backward looking
1/9
What if natural language were not just a prompt, but the genome of an evolutionary system?
In “ELMER: Evolutionary Language Model that Explores and Refines,” Ahmed Khalifa, Julian Togelius, and I trained an 8B model to evolve executable trading policies in natural language. 🧵
Season 1.5 of Alpha Arena has officially ended !
- Mystery Model (a.k.a GROK 4.20) is the winner, up 12% on avg.
- Not only did it win, it made money in all four competitions
- GPT5.1 🥈 came in 2nd, and Gemini 3 🥉 3rd
- All trades & model outputs are 100% verifiable 👇
Modern LLMs are good at writing code, but not necessarily at optimizing policy. But you can wrap the LLM in an evolutionary framework, and search for policies. In a new @the_nof1 paper, we show that we can automatically generate strong trading policies.
Season 1.5 of Alpha Arena is now LIVE with $320K deployed
It features:
- Multiple competitions
- Tons of new data
- 2 new models
- US equities
Most AI benchmarks test knowledge, our goal is to test judgement
Watch live below 👇
The next season of our benchmark will have lots of improvements. Also, we have plenty of other things going on at @the_nof1 which we haven't made public yet. Markets are fun to play, and make AI players for.
Qwen's portfolio is up +60%
Gemini's is down -60%
Of course, too early to tell how much is skill vs. noise
Next season we'll run many instances of the models in parallel for statistical rigor
The goal of Season 1 was to look for biases. What are the major differences between the LLM's trading styles, even with the same prompt? Can they even follow basic risk management rules?
A few early patterns:
> Qwen has only made 22 trades. It almost *never* has more than two positions on
> Gemini has made 108 trades. It literally always has the max number of positions on (6)
> Qwen has higher self-reported confidence (avg. 80% vs 65%)
> Qwen's stop loss and take profit levels are *much* tighter than Gemini's, but Gemini breaks its own rules often, and gets out early (others don't do this)
Overall, we're excited by the potential of LLMs and trading, but we're still skeptical. Much to test and learn
So much of finance happens behind closed doors. As you know, it's an insanely secretive industry
We're excited to put more experiments out in the open, both in trading and AI research
The next season of Alpha Arena will include a human trader, as well as our homegrown models
The only thing I accept being harder than @NetHack_LE are benchmarks on live real-world data. Well done @jay_azhang & @togelius from @the_nof1! Excited to see how AI will perform over longer investment timelines.