A team of PhD and Masterโs students obsessed with AI Agents. ๐ค
We research AI agents that can do everything โ and ones that do one thing really, really well.๐ฅฐ
Meet Scale-SWE: the largest open-source SWE dataset with real tasks.
With part of Scale-SWE, Qwen3-30A3B reach 64% on SWE-bench(V), surpassing GLM-4.7-Flash (59.2%).
20k SWE instances and 71k trajectories already public. Much more data on the way!
https://t.co/LaGdzfZ3Wk
๐ Introducing BeyondSWE: a benchmark for AI agents to tackle real-world engineering beyond single-repo fixes.
We evaluate Cross-Repo, DomainFix, DepMigrate, and Doc2Repo tasks to study the fusion of DeepResearch and coding.
๐ท https://t.co/VPngWRo7PX