Happy to share that MURPHY has been accepted at #NeurIPS2026! 🎉
As we move toward more agentic training and iterative self-improvement, models need to learn from the full process of trying, failing, receiving feedback, and correcting themselves, not just from the final outcome. 🔁🧠
MURPHY makes GRPO feedback-aware, assigning credit to earlier attempts when their feedback helps enable a later successful fix. 🛠️✅
🎥 Quick overview below
📄 Paper: https://t.co/W6cDZFoxcA
Can we predict how agent swarms will behave without burning thousands of dollars💰 on experiments?
We built a simulator that generates task DAGs (from real RSI experiments) and lets swarms re-play with them: a cheap way to study swarm scaling before paying for the real thing.
Update your agent just by talking to it!
Reef introduces the /reefine command: just describe what you want your agent to do. Reef builds the change, checks it, and ships a new version of the agent with an updated harness.
Examples:
💬 "> /reefine add a /chat mode for faster responses"
🔊 "> /reefine tell me out loud when you're done"
⚡ "> /reefine add jev as a tool": 120 papers screened in 9 s
Quick and personalized harness evolution will be key for continual learning and self-improving agents.
Reef also supports many other functions, including test-time training and model evolution. Try it here: https://t.co/EMz1yVcHVf
Excited to share what we have been building at @MIT
A new way to evaluate social simulation: Social Simulation vs. The Real Future. Each week, simulators predict how people will think, act, and change. Predictions are locked in before the results come out. We then score them and publish the scores.
🏆 Building a simulator? Bring it: https://t.co/iuSvSlbtZh
🔑 Want to build the arena with us? It is open source: https://t.co/pjwtIeRQpQ
Reef just dropped v0.1.1 🚀
In the past week, Reef has:
→ 🌟 Crossed 5K GitHub stars, adding ~1.6K
→ 🔧 Contributions from 35 developers so far
Try out v0.1.1 and tell us what you want your agent to keep learing.
Checkout Github Repo at: https://t.co/84OF3sXMxu
Reef just dropped v0.1.1 🚀
In the past week, Reef has:
→ 🌟 Crossed 5K GitHub stars, adding ~1.6K
→ 🔧 Merged 46 PRs, with contributions from 12 developers
Try out v0.1.1 and tell us what you want your agent to keep learing.
Checkout Github Repo at: https://t.co/84OF3sXMxu
Your agent can now grow new abilities, just by asking.
Reef now supports personalized harness evolution with /reefine: describe what you want your agent to do. Reef builds the change, checks it, and ships a new version.
Examples:
💬 "> /reefine add a /chat mode for faster responses"
🔊 "> /reefine tell me out loud when you're done"
⚡ "> /reefine add jev as a tool": 120 papers screened in 9 s
Works with any agent harness through a Reef adapter: Codex, Hermes, OpenCode, Pi and Terminus 2 today.
Try it here: https://t.co/yUfl1eoj6F
#agent #harness #rsi #llm
🔥 Jev is on fire!
😎 So we ask Reef: /reefine add jev as a tool
Reef then evolves the agent, and adds Jev into its harness.
We use this evolved agent to hunt for papers related to "Self-evolving Agents" published in 2026 on arXiv. The agent found and screened 120 papers in 9 seconds!
Come and use Reef to evolve your agent: https://t.co/IsAcJZFNQq
Really happy to see FrontierSmith accepted to NeurIPS as a Spotlight!
It’s been a great year working with @wenhaocha1 , and I’m very proud of what the FrontierCS team has built together. We also have more and more new work coming soon. Excited to share them!
Introducing HyperTrace (Hypothesis-Based Preference Tracing).
We formulate online LLM personalization as latent preference tracing: maintaining natural-language hypotheses about user preferences and continually updating them with SMC-style reweighting.
Full thread 👇
Reef's new /reefine command helps you equip your harness with new capabilities, not just micro-optimizations that gets you higher on benchmarks.
Checkout the demo video for the example cases!
Code is open-sourced at https://t.co/84OF3sXMxu
🚀 Reef now supports GEPA for prompt optimization and Meta-Harness for harness optimization!
Both run on Reef’s own harness evolution backend and work across a variety of harnesses (you can integrate your own by adding an adapter). The backend maintains an algorithm state that tracks past attempts, their relationships, and their performance.
Bring your tasks and an evaluator. Let Reef evolve your prompts or your entire harness!
Check it out here: https://t.co/IOhoLExXmI
Thanks @simon_ycl for implementing and maintaining this feature!
Don’t have GPUs? No problem.
Reef can now train through the Tinker API. Just change one line in your config and turn production feedback into continuously improving models.
🚀 Reef now supports training with the Tinker API!
Unlock continual learning with LoRA training by changing just one line in your config. Collect feedback, train, and publish versioned updates. No local GPU required.
Thanks to @Jayzou3773 for implementing this feature! 🙌
Check out our repo and contributions are always welcome! 👇
https://t.co/IOhoLExXmI
🧵 With unlimited compute, how fast can agents surpass humans? We introduce Elo-per-token analysis to profile agent performance curves across multiple open-ended tasks.
• Agents initially scale faster than repeated sampling, but over long horizons converge toward their theoretical log-linear scaling curve. Humans, in contrast, improve superlinearly.
• These curves also tell us how to spend test-time compute: the scaling inflection point gives a simple rule for splitting a fixed budget across agent sessions. Split a long run in a principled way, and you can get significant gains over a single run.
• Fitting human-time and agent-token curves also gives us a fun way to translate AI compute into human time. Taking OpenAI’s ~130B-token Navier–Stokes run as input and extrapolating across the two curves gives an equivalent of ~41 years of work by a mathematician at 8 hours/day 😮.
we have integrated bionemo inference runtime to speed up our inference.
great to see our models with optimized gpu kernels for faster and more efficient biomolecular structure prediction.
faster inference → faster iteration → faster science. 🚀