Scientific agents should improve through real research—with researchers helping shape what “better” means.
We’re glad to work with @PhAILabs as a joint R&D and technical support partner on ScienceBuddy.
For us, the exciting question is how to connect research tasks, tool interactions and researcher feedback into a process that can be evaluated and improved over time. This is where data, evaluation and engineering come together.
Congratulations to the PhAI Labs team on the launch. We look forward to learning alongside the researchers who put ScienceBuddy to work.
Today, PhAI Labs launches ScienceBuddy, an interactive research workspace for scientific agents that improve through researcher collaboration.
Initiated by PhAI Labs, with @muchencq (Muchen AI) as a joint R&D and technical support partner. 🧵
ScienceBuddy couples two loops:
🔹 An inner loop that refines the agent harness—its instructions and skills.
🔹 An outer loop that trains the model through rubric-guided reinforcement learning.
Together, they form Recursive-in-Recursive Self-Improvement.
On a held-out set of 180 scientific problems, pass@4 coverage with Qwen3.5-4B rose from 48.3% to 67.8% under the same four-attempt budget.
DFM sets the direction—from predefined tasks toward open-ended discovery. ScienceBuddy explores how scientific agents can improve through sustained collaboration with researchers.
📄 Tech report: https://t.co/gxPRnp5aZJ
💻 GitHub: https://t.co/xNlsshOyXs
🔬 Product access: https://t.co/ksVs3zUXd9
#PhAILabs #DFM #ScienceBuddy
Hello, X. We’re Muchen AI (沐晨科技), a data intelligence services company based in Chongqing, China.
Our foundation is AI data delivery. We’re building on that foundation in two directions: defining and evaluating model capabilities, and helping enterprises put data intelligence to work.
Our focus is Long-horizon AI Data Operations: connecting task definition, data engineering, evaluation, quality governance and feedback across repeated cycles of improvement.
We believe AI’s value should be judged by what it can reliably do in the real world. That takes clear tasks, useful data and evaluations that reveal where systems still fail.
Here, we’ll share practical notes on AI data, model and agent evaluation, and lessons from building AI-enabled workflows.
Building or evaluating AI? We’d love to compare notes.
Data Intelligence. Real-World AI Value.
https://t.co/v3jIQZERDJ