@arjunomics Crazy to see Claude Code think 3x as hard for the same average performance.
> The same Opus 5 under Pi costs about a quarter as much for the same accuracy
@badlogicgames check this out!
On our benchmarks, Kimi thinks it's being evaluated. I wonder if it will still think it's being evaluated when real researchers are running real queries on something like OpenCode. Essentially, can it recognize questions and formats that look like benchmark runs instead of real usage, or does it always assume it's being tested?
I think this should inform future benchmark design.
@moinnadeem Modal is expensive for agents running computationally heavy science work loads I.e. tens of thousands of agents that need 6 cpus and 32 GiB - compute requirements only increase.
For biotech, at least now, Modal isn't feasible and you have to build in house alternatives.
I published a new blog today!
It is widely known that prompting can have small effects on performance, but in this blog we show 2 lines of behavioral prompting can substantially impact a model’s ability on frontier biology tasks. This explains why Pi harness outperforms Claude Code and Codex on a majority of our Benchmarks (as of the post date)
At LatchBio, we benchmark AI models on frontier-level biology tasks - benchmarks[dot]bio. This blog explores Tx-Bench-PP and ScBench-Long.