our harness + Opus on MLE bench lite also outperforms other harnesses
MLE bench is the closest public benchmark to our research tasks, and we chose MLE lite because full MLE bench runs just take too long
is the long runtime of ML experiments the reason why there are so few such benchmarks?
our general research harness defines a new Pareto frontier on terminal bench 2.1
improving general research skills seems to help with SWE tasks too
here is how our harness works:
- async orchestration layer
- extremely parallelised execution runtime
- dedicated communication system between main/ subagents
- native research support: efficient tracking, identification & querying of experiments
and the beautiful office was provided to us kindly by fellow YC company https://t.co/JZOKgdQsqN - they've just raised $45m to build autonomous white collar workers
yesterday we hosted an incredible hack for RSI & autoresearch in SF with Autolab (@ottogin1 and the group are absolute pros at organising hackathons)
we had almost 400 registrations and sadly could not take in everyone, but stay tuned for future hacks.
teaser: there might be one just around the corner next Sunday on #RSI & Harnesses
inspecting the traces shows that the necessary environment completion rule 'is_done=True' gets lost in the model compaction.
Out of all Sol runs, only 29% final compacted context still contains this rule for high and xhigh. Compared to 51% at low reasoning
SOTA models reach <30% on our long-horizon benchmark.
we built 100 tasks where models operate a multi tool trading system and choose/ tune available tools
turns out higher reasoning confuses Sol and it fails to submit the necessary final order, hence worse rewards
Today we are launching EdotEnv, a Quant Neolab building toward RSI.
RSI needs a loop of increasingly difficult tasks, which markets naturally are: Trading well means markets become more efficient, this makes successful trading harder. Reach out if you are interested!