Excited to share that our Enterprise AI team at @scale_AI@ScaleAILabs has two papers accepted to @icmlconf this year.
What makes this especially meaningful is the team (Andrew Klearman Varun Ursekar Apaar Shanker Radu Revutchi Veronica Chatrath Rohin Garg Rishav Chakravarti Sam Denton) behind it. Many of the authors are new and early in their research journey, while at the same time being deeply involved in real customer deployments. This dual lens -- building in production while pushing research forward -- is exactly where I believe the most important innovation is happening. Both works are grounded in a common motivation: agent engineering needs in real-world enterprise settings.
1. VeRO: An Evaluation Harness for Agents to Optimize Agents [https://t.co/uFdTTyjgqa] turns agent self-optimization into a systematic, measurable, and comparable process by introducing a reliable evaluation environment which ensures reproducible results and a benchmark with a standardized suite of target agents, tasks, and evaluation procedures.
2. Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation [https://t.co/l8QvkhgkQY] focuses on trustworthy evaluation. It highlights that the evaluation quality is fundamentally constrained by how evaluation sets are constructed. By introducing semantic stratification, it establishes a formal semantic coverage framework which enables robustness across retrieval regimes and provides interpretable visibility into failure modes.
Proud of the team for pushing on problems that sit at the intersection of research and real-world deployment. This is just the beginning.
#icml #enterpriseAI #reliableAI
Excited to share that our Enterprise AI team at @scale_AI@ScaleAILabs has two papers accepted to @icmlconf this year.
What makes this especially meaningful is the team (Andrew Klearman Varun Ursekar Apaar Shanker Radu Revutchi Veronica Chatrath Rohin Garg Rishav Chakravarti Sam Denton) behind it. Many of the authors are new and early in their research journey, while at the same time being deeply involved in real customer deployments. This dual lens -- building in production while pushing research forward -- is exactly where I believe the most important innovation is happening. Both works are grounded in a common motivation: agent engineering needs in real-world enterprise settings.
1. VeRO: An Evaluation Harness for Agents to Optimize Agents [https://t.co/uFdTTyjgqa] turns agent self-optimization into a systematic, measurable, and comparable process by introducing a reliable evaluation environment which ensures reproducible results and a benchmark with a standardized suite of target agents, tasks, and evaluation procedures.
2. Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation [https://t.co/l8QvkhgkQY] focuses on trustworthy evaluation. It highlights that the evaluation quality is fundamentally constrained by how evaluation sets are constructed. By introducing semantic stratification, it establishes a formal semantic coverage framework which enables robustness across retrieval regimes and provides interpretable visibility into failure modes.
Proud of the team for pushing on problems that sit at the intersection of research and real-world deployment. This is just the beginning.
#icml #enterpriseAI #reliableAI
Today, we're unveiling two new open-source AI robots! HopeJR for $3,000 & Reachy Mini for $300. DM me if you want to be added to the waitlist 🤖🤖🤖
Let's go open-source AI robotics!
Excited to present FastTD3: a simple, fast, and capable off-policy RL algorithm for humanoid control -- with an open-source code to run your own humanoid RL experiments in no time!
Thread below 🧵