Tomorrow in SF, we’re teaming up with Reinforce Labs to talk about what it takes to build AI agents people can trust.
Leaders from Google, Meta, Glean, Salesforce and Across AI join our @Centific_CAIR team on evals, guardrails and accountability in production.
Don’t forget to register if you haven’t yet. See you tomorrow!
https://t.co/N39nC4VbDw
#COLM2026
A robot fails. Should it retry, check a sensor, or ask a person? Today’s AI can’t reliably make that call.
So we built a benchmark where every failure’s true cause is known, then tested six vision-language models. Three findings: they diagnosed failures no better than guessing, their confidence meant nothing, and when we made asking 4x costlier they asked just as often.
The fix wasn’t a smarter model. It was the right sensor: grasp failures were nearly invisible to the camera but almost fully readable from the robot’s force data. And one question to a person made a model about as accurate as the human it asked.
Takeaway: don’t trust the model’s confidence. Measure its accuracy, show it its own sensor data first, and ask a person only when no sensor has the answer.
Centific’s paper by Eshika Pathak and Leela Krishna, accepted at the IEEE IROS 2026 Workshop on Human-Robot Dialogue. @krishnacentific is presenting it tomorrow in Pittsburgh. Come say Hi.
Read about the paper: https://t.co/wx4RfwowP3
#IROS2026 #Robotics #Centific
Launching Centific Physical AI, powered by HumanoidIQ, today at #IROS2026.
Training-ready robotics data: real demos, QC'd, replayed on a virtual robot, annotated step by step, standardized. Ready-to-license in our Data Marketplace.
At IROS? Let's talk. 🔗 https://t.co/O2KRzT6R2x
#PhysicalAI #Robotics
We’re excited to announce that two papers from the Centific Physical AI team have been accepted at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026 next week, both by Eshika Pathak and Krishna L.
"When Should a Failing Robot Ask?" at the IEEE Workshop on Human-Robot Dialogue, Thursday 1 October. https://t.co/wx4RfwowP3
"Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics" at the Workshop on Embodied Neuro-Symbolic AI for Reliable and Safe Robotics, Sunday 27 September. https://t.co/8kIw8J5a8K
Come and find us at either session. @ieeeiros
#IROS2026 #Robotics
I got to visit one of the best research paper reading groups in Bay Area - Saturday Robotics. @aurorafeng_01 - great summary of data pyramid for frontier robotics where often the industry gets lost in the weeds with VLA/WAM or egocentric vs teleoperation. Truth is its 80-20
🤖 On 8/20 at WRC Beijing, @saturdayrobotic partnered with EAI Time for our first event in China (and second event in Asia!). We invited 8 researchers and builders to experiment something different from our usual reading club: a live technical debate.
1. If you had $100M for embodied pre-training data, how would you spend it?
One proposal: 50% ego-centric data, 30% intermediate action data, 20% real-robot data. Another: go as far as 99:1, spending $99M collecting and annotating high-quality ego data and only $1M on edited data. First, human video is not tied to a particular robot embodiment. As hardware iterates, robot-specific data can depreciate, but same raw human video actually becomes more valuable. Second, there is a distribution problem with teleop. A teleoperator tends to repeat a small number of successful trajectories. Collecting another 10,000 repetitions doesn’t necessarily give pre-training the diversity it needs.
But there were strong arguments against spending $100M this way. One position favored pure action data: stronger action supervision may cost more to collect, but can reduce the GPU compute needed downstream. Another argued for starting with proprietary real-robot data in a vertical, accumulating actual deployment know-how first, and only scaling ego once collection costs fall. And one answer rejected the premise entirely: $100M isn’t enough to properly pre-train a foundation model anyway.
2. Does WAM actually have a higher ceiling than VLA?
Most agree that WAM does have a higher ceiling than VLA. It's because the learning paradigm is different: robotics doesn’t have enough high-quality labeled action data, so moving from VLA toward WAM resembles the broader shift from supervised → self-supervised learning. If WAM can learn from vastly more unlabeled video, its scalable data ceiling is much higher.
A counterargument is that given infinite high-quality data, WAM and VLA may have the same theoretical ceiling. WAM’s advantage today may simply be that its floor is higher under the data regime we actually have, particularly OOD.
A third position challenged the WAM/VLA distinction itself. The more fundamental question is how much high-order physical information a model can extract from data (geometry, dynamics, contact, state transitions, etc.) rather than what architectural label we put on it.
3. How should we evaluate embodied models without making iteration prohibitively expensive?
One proposal was that world-model/simulation-based evaluation could cover ~90% of generalization testing, with hardware reserved for final verification and high-precision tasks.
The interesting argument was that a simulator doesn’t necessarily need to reproduce reality perfectly to be useful. For development, what matters first is correlation.
But the boundary was clear: for high-precision manipulation, real-robot evaluation remains the only reliable final standard.
@aurorafeng_01@saturdayrobotic Thank you for sharing, absolutely agree its about data scale va precision in both case of ego vs teleop as well as wam vs vla
@LinfengZhaoZLF@saturdayrobotic Excellent the promise of robotics pipeline automation building closed loop action-perception feedback is an excellent choice for advancing robotic frontier
We are excited to release our work on DexPilot, a markerless, glove-free and vision based teleoperation of dexterous robot hand-arm system
pdf is here https://t.co/qWD8KURCum
link to more videos https://t.co/dFoQGtWaOP
https://t.co/L8g8Mqsp2k
We're excited to announce that 4 Centific papers are being presented at KDD 2026 workshops in Jeju, South Korea.
📍 Enterprise AI Agents: From Prototypes to Production (Aug 10, 8am–12pm, ICC Jeju 402A)
1️⃣ RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents
Dialogue agents are trained on cooperative users and deployed against hostile ones. RL-ADA hardens agents with world feedback, no human labels in the loop.
2️⃣ Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows 📄 https://t.co/OP5qMHHAHb
Five environments at Jira REST v3 / Confluence v2 schema fidelity, rewards computed entirely from the tool-call trace: no live API, no learned judge, no human label. GRPO lifts average reward from 0.35–0.92 to 0.95–1.00 on a 4B model.
3️⃣ Rethinking MedAgentBench: A Framework for Fair Medical LLM Agent Evaluation
Four evaluation failures let a do-nothing agent score 42%. Chief among them: branch imbalance, with four v2 task types placing 70–97% of instances on the no-action branch. We rebuild it as MAB-v3.
📍 RelSciFM: Reliable Scientific Foundation Models (Aug 10, 1–5pm, ICC Jeju 402A)
4️⃣ BCoughBench: Benchmarking Respiratory Acoustic Foundation Models Under Body-Coupled Wearable Sensor Conditions 📄 https://t.co/nbAYxVbAQV
Respiratory FMs are benchmarked on smartphone audio; deployment is moving to wearables that attenuate high frequencies through tissue and bone. Across 5 FMs and 9 tasks, mean AUROC falls from 0.785 to 0.689–0.723, and no model clears the clinical sensitivity bar (Se@Sp95 ≥ 0.20) on any disease task. Sex classification collapses (0.954 → 0.60); COVID detection is untouched (Δ = −0.004).
Common thread: models look good in test but fail in production, because the test environments aren't a true representation of production. We swapped the sensor (BCoughBench), the user (RL-ADA), the workflow (RLVR), and the scoring rules (MedAgentBench), and the numbers move.
If you're at KDD in Jeju, come say hello 👋
#KDD2026 #EnterpriseAI #AgenticRL #AIAgents #FoundationModels
Pretty slick edge model from #nvidia.
#centific pilots in warehouses, drones, mobile inspection & sensor-triggered reasoning. Frame answers in battery, bandwidth constrained devices fusing cameras with IMUs, a motion signature, fall, collision and abrupt maneuvers.
Introducing Cosmos 3 Edge: our open frontier world model built to run on-device.
Cosmos 3 Edge helps robots learn and act, autonomous vehicles understand road scenes and predict intent, and vision AI agents reason across live video for smart infrastructure.
With 4B parameters and a 2B Nemotron-based reasoner, you can run it on DGX Spark, NVIDIA Jetson, and more.