1121 cycles.
New record on Anthropic’s take-home kernel challenge.
One meta-evolve agent. 200 attempts.
Stop only evolving solutions.
Evolve evolution.
Introducing Atoms: the first AI team that builds real businesses.
From research to build, launch, and scale, all autonomous.
Don’t Vibe Code. Vibe Business.
→ https://t.co/g0L4NRU7xE
Would anyone be interested in scaling completely different Agent environments? 👀
Teaser: We're open-sourcing AutoEnv — generate a full env (text, multimodal, even 3D) + scripts to continuously scale in-world data for just ~$4 💸
Textual scaling is available at: https://t.co/FPJvtyPE7B
See Generated Multimodal Env Below:
Is ReAct the final for LLM Agents? 🤔
Not anymore. ReAct is stuck step-by-step. Planning agents separate plans from actions.
Introducing ReCode: agents control their decision granularity, like humans do.
20.9% better, 78.9% cheaper, 3.7× data efficient. The ReAct era is over.
1/5
🤖Check The Hitchhiker’s Guide to Agents HERE🤖
Our Foundation Agents Survey V2 level up to 396 pages – every chapter is a full-on survey itself!
🧠 Agent Framework & Components
🌍 World Model & Memory
🔄 Self-Evolution
👥 Multi Agents
🛡️ Safety
1/4
It's actually a pity that we got no enough time to maintain OpenManus during the past 3 months.
But the better news is that we will build a formal open-source community for OpenManus at the end of this month.
🧠264 pages and 1416 references chart the future of Foundation Agents.
Our latest survey dives deep into agents—covering brain-inspired cognition, self-evolution, multi-agents, and AI safety.
Discover the #1 Paper of the Day on Hugging Face👇:
https://t.co/CvtqGKyCYb
1/3
Text-to-SQL woes? Reasoning models stumble in zero-shot tasks 😓
Enter Alpha-SQL — our breakthrough boosts 7B LLMs by 15-20%, topping GPT-4o SOTA and even reasoning models on BIRD! 🎉
Test Time Scaling still shines.
How we nailed it 👇:
1/5