During COLM week, @Centific and Reinforce Labs are bringing researchers, builders, and product leaders into one room to talk about what it takes to build agents people can trust.
Fireside: Vinay Rao (VP, @Google) ร Anish Das Sarma (CEO, Reinforce Labs)
Panel:
- Nilou Salehi: CEO & Co-founder, Across AI
- Panagiotis Papadimitriou: Sr. Director, @Meta
- Nilesh Dalvi: Engineering Leadership, @glean
- Manjeet Singh: Sr. Director, @salesforce
- Moderated by Prasanna Desikan, Head of Healthcare AI, Centific
Evaluation before deployment, guardrails that hold in production, accountability when an agent goes off-script, and what it takes to earn trust.
๐ Oct 8, 2026 ยท 6:00โ9:00 PM
๐ San Francisco (address after you register)
Capped at 50. RSVP early: https://t.co/SRFIDT2c4Q
#COLM2026
During COLM week, @Centific and Reinforce Labs are bringing researchers, builders, and product leaders into one room to talk about what it takes to build agents people can trust.
Fireside: Vinay Rao (VP, @Google) ร Anish Das Sarma (CEO, Reinforce Labs)
Panel:
- Nilou Salehi: CEO & Co-founder, Across AI
- Panagiotis Papadimitriou: Sr. Director, @Meta
- Nilesh Dalvi: Engineering Leadership, @glean
- Manjeet Singh: Sr. Director, @salesforce
- Moderated by Prasanna Desikan, Head of Healthcare AI, Centific
Evaluation before deployment, guardrails that hold in production, accountability when an agent goes off-script, and what it takes to earn trust.
๐ Oct 8, 2026 ยท 6:00โ9:00 PM
๐ San Francisco (address after you register)
Capped at 50. RSVP early: https://t.co/SRFIDT2c4Q
#COLM2026
Opus 5.5 ranks first on our enterprise sales benchmark, and it's strongest where the work is hardest: Revenue Ops.
It's also unusually reliable. Across 330 runs, it went in circles once and never left an error unhandled.
10 models. 33 tasks. Replicas of enterprise applications.
A ranking hides where a model is actually strong. We report by role and by task.
Benchmark your own enterprise workflows:
https://t.co/VYs2dSZYMx
Agent-to-agent marketplaces now let autonomous agents negotiate, transact, and settle payments on our behalf, with no human in the loop. That autonomy makes what the agents actually do inside the marketplace worth understanding.
A new @Centific paper at the #COLM2026 workshop on Agent Behavior argues these marketplaces should be evaluated as behavioral systems, not scoreboards.
Across 140 rollouts in two marketplaces, one priced (MarketDeal) and one barter-only (SwapShop), the close rate barely moved. Almost everything underneath it did.
A close rate can't tell you:
โ whether both sides gained, or one took the whole surplus
โ what the agent checked, revealed, and where the money went
โ whether it can trade at all without a price
Two models, same family, same tools, behaved nothing alike. Behavior depended on the market's structure, not just the model's capability.
Centific's rubric scores seven dimensions of agent behavior from the trajectory, not the ledger.
๐ Paper by @shivalidalmia10, @jainashi03, and @mukherjiab: https://t.co/UG4QMcHJCK
โโโโโโโโโโโโ
Please also join us
๐ Oct 8 in SF: Agents You Can Trust, A COLM Happy Hour with Vinay Rao (@Google), @AnishDasSarma (Reinforce Labs), @nilou_salehi (Across AI, @UCBerkeley), @ppapadim (@Meta), Nilesh Dalvi (@glean), @CoachManjeet (@salesforce ), @DesikanPrasanna (Centific)
RSVP: https://t.co/kIFGeRVdqu
How do you know if a robot picked the right stored demonstration for a new scene?
Normally you only see whether its choice worked, never how the alternatives would have done. So we executed every stored experience in every query scene and built the complete table of outcomes.
The result: visual similarity predicts whether a given demonstration-scene pair will succeed very well (AUROC up to 0.96) and ranks the candidates within a single scene no better than chance.
Predicting transfer and selecting an experience are different problems. To learn more, read the full paper by Eshika Pathak & @krishnacentific , accepted at the #IROS2026 Workshop on Embodied Neuro-Symbolic AI for Reliable and Safe Robotics: https://t.co/8kIw8J5a8K
#Robotics #Centific
Launching Centific Physical AI, powered by HumanoidIQ, today at #IROS2026.
Training-ready robotics data: real demos, QC'd, replayed on a virtual robot, annotated step by step, standardized. Ready-to-license in our Data Marketplace.
At IROS? Let's talk. ๐ https://t.co/O2KRzT6R2x
#PhysicalAI #Robotics
A robot fails. Should it retry, check a sensor, or ask a person? Todayโs AI canโt reliably make that call.
So we built a benchmark where every failureโs true cause is known, then tested six vision-language models. Three findings: they diagnosed failures no better than guessing, their confidence meant nothing, and when we made asking 4x costlier they asked just as often.
The fix wasnโt a smarter model. It was the right sensor: grasp failures were nearly invisible to the camera but almost fully readable from the robotโs force data. And one question to a person made a model about as accurate as the human it asked.
Takeaway: donโt trust the modelโs confidence. Measure its accuracy, show it its own sensor data first, and ask a person only when no sensor has the answer.
Centificโs paper by Eshika Pathak and Leela Krishna, accepted at the IEEE IROS 2026 Workshop on Human-Robot Dialogue. @krishnacentific is presenting it tomorrow in Pittsburgh. Come say Hi.
Read about the paper: https://t.co/wx4RfwowP3
#IROS2026 #Robotics #Centific
Weโre excited to announce that two papers from the Centific Physical AI team have been accepted at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026 next week, both by Eshika Pathak and Krishna L.
"When Should a Failing Robot Ask?" at the IEEE Workshop on Human-Robot Dialogue, Thursday 1 October. https://t.co/wx4RfwowP3
"Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics" at the Workshop on Embodied Neuro-Symbolic AI for Reliable and Safe Robotics, Sunday 27 September. https://t.co/8kIw8J5a8K
Come and find us at either session. @ieeeiros
#IROS2026 #Robotics
Qwen 3.8 Max completes an enterprise sales task on the first try more often than any of the seven models we tested: 13.9%, against 12.4% for the leaderboard leader.
These are long-horizon tasks across eight or more enterprise applications. No model completed more than 27% even with five attempts, and Qwen sits 4th on that five-attempt board. A deployed agent gets one attempt.
It's one real weakness is navigation: lost between applications in 21% of its failures, against 7% for the others. That is a behaviour, and behaviours can be trained.
[email protected]
It completed 3 tasks no other model could. It also finished without checking its own output in 10% of its failures, against 4% for the others.
Incomplete briefs are where enterprise agents break.
We grade for it both ways.
https://t.co/ya6zeZpXPc
Muse 1.3 is the strongest model in our enterprise sales benchmark when it has the full brief, and it gives up more than any other model once details go missing.
All required task details provided: 80.0 out of 100, best of the seven models we tested. Key details omitted: 53.6, and 4th. A 26-point fall, larger than any other model's.
๐ ๐๐ฒ๐ป๐๐ถ๐ณ๐ถ๐ฐ ๐ถ๐ ๐ฒ๐ ๐ฐ๐ถ๐๐ฒ๐ฑ ๐ณ๐ผ๐ฟ ๐๐ต๐ฒ ๐ฟ๐ฒ๐ฐ๐ฒ๐ป๐ ๐ข๐ฝ๐ฒ๐ป๐๐ ๐๐๐๐ฟ๐ฎ ๐ฎ๐ป๐ฑ ๐๐ป๐๐ต๐ฟ๐ผ๐ฝ๐ถ๐ฐ ๐๐ฎ๐ฏ๐น๐ฒ ๐ฑ.๐ญ ๐น๐ฎ๐๐ป๐ฐ๐ต๐ฒ๐.
So we put ๐๐ฃ๐ง-๐ฒ ๐๐๐๐ฟ๐ฎ ๐ฃ๐ฟ๐ผ and ๐๐ฎ๐ฏ๐น๐ฒ ๐ฑ.๐ญ, plus 5 other frontier models, to test.
โ๏ธ ๐ฒ๐ป๐๐ถ๐ฟ๐ผ๐ป๐บ๐ฒ๐ป๐
Centific's order to cash RL environment, 8+ applications and connectors (D365, Workday, chats, emails, SharePoint Docs) with 200+ tools.
๐ฏ ๐ง๐ต๐ฒ ๐๐ฎ๐๐ธ๐
33 long horizon sales and revenue workflows that mirror what BDR managers, quota carrying sellers and RevOps teams do every day. Each takes a professional 5+ hours to several days. Model runs span hundreds of tool calls.
๐ ๐ช๐ต๐ฒ๐ป ๐ฎ๐น๐น ๐ฟ๐ฒ๐พ๐๐ถ๐ฟ๐ฒ๐ฑ ๐๐ฎ๐๐ธ ๐ฑ๐ฒ๐๐ฎ๐ถ๐น๐ ๐ฎ๐ฟ๐ฒ ๐ฝ๐ฟ๐ผ๐๐ถ๐ฑ๐ฒ๐ฑ
๐ฅ GPT-6 Astra Pro: 13/33 completed
๐ฅ Fable 5.1: 10/33 completed
๐ ๐ช๐ต๐ฒ๐ป ๐๐ฒ ๐๐ฒ๐น๐ฒ๐ฐ๐๐ถ๐๐ฒ๐น๐ ๐ผ๐บ๐ถ๐ ๐ธ๐ฒ๐ ๐๐ฎ๐๐ธ ๐ฑ๐ฒ๐๐ฎ๐ถ๐น๐
๐ฅ GPT-6 Astra Pro: 5/33 completed
๐ฅ Fable 5.1: 2/33 completed
Accuracy drops sharply. Incomplete briefs are where enterprise agents break.
๐ ๐ง๐ต๐ฟ๐ฒ๐ฒ ๐ผ๐ฏ๐๐ฒ๐ฟ๐๐ฎ๐๐ถ๐ผ๐ป๐ ๐๐๐ผ๐ผ๐ฑ ๐ผ๐๐
๐ง ๐ญ๐ฌ ๐๐ผ๐ฟ๐ธ๐ณ๐น๐ผ๐๐ ๐ฟ๐ฒ๐บ๐ฎ๐ถ๐ป ๐ผ๐ฝ๐ฒ๐ป ๐๐ฒ๐ฟ๐ฟ๐ถ๐๐ผ๐ฟ๐. No model completed them. Astra came closer than any other.
๐ง ๐ฃ๐น๐ฎ๐ป๐ป๐ถ๐ป๐ด ๐ถ๐ ๐ฎ ๐ฐ๐น๐ฒ๐ฎ๐ฟ ๐๐๐ฟ๐ฒ๐ป๐ด๐๐ต ๐ณ๐ผ๐ฟ ๐๐๐๐ฟ๐ฎ. Only 3 of its 48 incomplete runs traced back to planning.
๐ ๐๐๐๐ฟ๐ฎ ๐ฎ๐ป๐ฑ ๐๐ฎ๐ฏ๐น๐ฒ ๐ผ๐๐ป๐ฒ๐ฑ ๐๐ต๐ฒ ๐ถ๐ป๐ฐ๐ผ๐บ๐ฝ๐น๐ฒ๐๐ฒ ๐ฏ๐ฟ๐ถ๐ฒ๐ณ ๐๐ฒ๐๐. Every other model completed zero.
โ Centific also provides environment audits and independent evals.
๐ฉ ๐ช๐ฎ๐ป๐ ๐๐ผ ๐น๐ฒ๐ฎ๐ฟ๐ป ๐บ๐ผ๐ฟ๐ฒ? DM me or reach out to [email protected]
#EnterpriseAI #ModelEvaluation #AgenticRL #Centific #Astra #Fable5point1
Fair read of the headline - but the zero-call agent wasn't efficient, it did nothing. No data touched, no audit run, one paragraph asserting the job was done. It only "aligned with expectations" because the rubric graded the prose instead of the trace. Efficiency and doing nothing look identical to that grader.
We gave an AI agent a pricing audit to run. It made 75 tool calls doing the work.
A second agent made zero calls. It wrote one paragraph saying the job was done.
The grader scored the second one higher.
That is not a bad model. That is a broken test.
@ToniRMCardoso Exactly the failure. Our graders were reading the summary, so a confident closing paragraph beat 75 real tool calls. We rebuilt them to score the trace: which tools fired, in what order, against what state. "Evidence over vibes" is the whole fix.
It was not a one-off. Across 16 tasks, the agent that did nothing beat the agent that did the work on 13 of them.
We found six failures like this. Here's what each one was and how we caught it:
https://t.co/qBciJto6xR
The first one is live. A sales operation across 22 connected systems, first lead to paid invoice. Open it on Hugging Face and run it yourself.
Browse all 24 โ https://t.co/ya6zeZppZE
Most RL environments simulate one or two pieces of real work. Ours are built from all of it.
Centific opened a catalog where AI agents have to do the job, not describe it. 24 environments: closing a deal, hiring someone, paying a supplier, shipping code, auditing a discharge summary.