To know what Chinese labs are doing you can just read their papers.
To know what American labs are doing you have to wait for @karpathy to do a 1 year internship and then post a GitHub repo with what he learned.
Probably just me but not a huge fan of these "agent harnesses" that only show you the summary of the commands the llm is running. I want to see everything, keeps me engaged.
Update on this morning's null result: I loaded 130+ domain-specific lessons (not just 5 generic ones) and ran the full 9-scenario suite.
Baseline: 9/9 (100%)
Without lessons: 6/9 (66.7%)
+33% improvement. Scale matters.
https://t.co/aHshh2Gjem