We are excited to announce a strategic partnership between Turing and @HUMAIN to build the world’s first enterprise-scale AI Agent Marketplace on HUMAIN ONE.
This collaboration brings together HUMAIN’s AI operating system and infrastructure with Turing’s expertise in frontier AI systems, evaluation, and deployment to unlock a new era of enterprise intelligence.
The HUMAIN ONE AI Agent Marketplace will enable organizations to:
-Discover and deploy AI agents across every business function
-Scale intelligent workflows across HR, finance, legal, operations, and beyond
-Build and monetize enterprise-ready AI agents in a secure, governed environment
“Superintelligence should not remain abstract. It should deliver productivity, increase ease of use, and unleash humanity’s untapped potential.”
— @jonsid, CEO and Co-Founder of Turing
Together, we are accelerating the shift from traditional software to agent-driven organizations, where AI not only supports work but executes it.
This partnership also marks an important milestone in advancing superintelligence from concept to real-world impact, across the Kingdom of Saudi Arabia and globally.
By combining advanced AI systems with human judgment and expertise, Turing and HUMAIN aim to unlock new levels of productivity, accelerate innovation, and drive long-term economic growth.
Learn more about how we are shaping the next generation of AI infrastructure and innovation below.
@FastCompany's World's Most Innovative Companies list is out — and Menlo's portfolio shows up strong. From @AnthropicAI to @UnstructuredIO, these are category leaders building companies that last. https://t.co/vRfUwxhVSP
Turing has been named one of @FastCompany's Most Innovative Artificial Intelligence Companies of 2026!
The recognition comes at a defining moment for AI.
Bigger models. More data. Greater compute. Now paired with AI coding tools that are helping build the next generation of systems.
Proud to be shaping what comes next!
Most AI agents today are failing the enterprise 'vibe check.' ServiceNow Research just released EnterpriseOps-Gym, and it’s a massive reality check for anyone expecting autonomous agents to take over IT and HR tomorrow.
We’re moving past simple benchmarks. This is a containerized sandbox with 164 database tables and 512 functional tools. It’s designed to see if agents can actually handle long-horizon planning amidst persistent state changes and strict access protocols.
The Brutal Numbers:
→ Claude Opus 4.5 (the top performer) only achieved a 37.4% success rate.
→ Gemini-3-Flash followed at 31.9%.
→ DeepSeek-V3.2 (High) leads the open-source pack at 24.5%.
Why the low scores? The research study found that strategic reasoning, not tool invocation, is the primary bottleneck. When the research team provided agents with a human-authored plan, performance jumped by 14-35 percentage points.
Strikingly, with a good plan, tiny models like Qwen3-4B actually become competitive with the giants.
The TL;DR for AI Devs:
✅ Planning > Scale: We can’t just scale our way to reliability; we need better constraint-aware plan generation.
✅ MAS isn't a Silver Bullet: Decomposing tasks into subtasks often regressed performance because it broke sequential state dependencies.
✅ Sandbox Everything: If you aren't testing your agents in stateful environments, you aren't testing them for the real world.
Read our full analysis here: https://t.co/CoONW0g9Pl
Check out the benchmark: https://t.co/BM5lSHzWX0
Paper: https://t.co/nE0j3QECMo
Codes: https://t.co/5Zj8WyZcI2
@ServiceNow@ServiceNowRSRCH@RajeswarSai@ShivaMalay@PShravannayak@TheJishnuNair@sagardavasam@SathwikTejaswi@tscholak@NVIDIAAI@turingcom@ServiceNowNews@jonsid
Turing contributed to Enterprise Ops Gym, ServiceNow’s new enterprise agent benchmark.
We designed 1,000 prompts across 8 enterprise scenarios spanning HR, CSM, ITSM, Email, Calendar, Drive, Teams, and hybrid cross domain workflows.
Tasks range from 7 to 30 steps with stateful system updates, expert reference traces, and deterministic verification scripts to evaluate correctness and policy compliance.
A rigorous step forward for long horizon enterprise agent evaluation.
Learn more:
Paper: https://t.co/JULVtkQLUK
Data Set: https://t.co/mGAlgvIiwM
Website: https://t.co/N56iNyPvkN
Code: https://t.co/DBC3SsvK6W
Our results point to three concrete research directions:
🗺️ Constraint-aware plan generation. Methods that reason over prerequisite structures before committing to actions.
🧠 Long-horizon state management. Episodic memory or structured state representations to curb error accumulation
🚦 Safe refusal and escalation. Agents must reliably detect infeasible or policy-violating requests and abstain cleanly.
We’re releasing 60% public version of the benchmark today. Come build on it ..we’d love to see what the community does with it. 🙏
By @ShivaMalay, @PShravannayak , @TheJishnuNair, @sagardavasam, @SathwikTejaswi , @tscholak@NVIDIAAI@turingcom@ServiceNowNews , @jonsid @Mila · With thanks to our partners at @Turing 🙏
🧵 Introducing 𝐄𝐧𝐭𝐞𝐫𝐩𝐫𝐢𝐬𝐞𝐎𝐩𝐬-𝐆𝐲𝐦🚀 : a rigorous new benchmark for stateful agentic planning and tool use in real enterprise environments.
1,150 expert-curated tasks · 512 tools · 164 DB tables · 8 domains. Every task verified by hand-written SQL, checking goal completion, state integrity and policy compliance🔥
𝐓𝐡𝐞 𝐡𝐞𝐚𝐝𝐥𝐢𝐧𝐞: Claude Opus 4.5 — our best-performing model succeeds on just 37.4% of tasks. With oracle tool access. No tool discovery required.
📄 https://t.co/S4EHVbabpy (trending #4 on daily-papers)
🌐 https://t.co/oMqtV1gHEG
🤗 https://t.co/Uv0pzQiuh2
💻 https://t.co/rohdFTfchK
🧵 Introducing ENTERPRISEOPS-GYM — a new stateful benchmark for enterprise AI agents.
Best model: 37.4% success — with oracle tool access. Give a tiny model a human plan? It catches up to giants.
The bottleneck isn’t tools. It’s planning. 👇
OpenEnv treats environments as first class infrastructure.
Through our collaboration with @AIatMeta and @huggingface, Turing is helping labs run tool-using agents against RL Environments that share:
-A standard step and reset API
-WebSocket sessions with per-client state
-MCP-style tool discovery and calling
-Observability hooks for rewards, errors, and drift
The payoff is simple. You can reuse the same evaluation pattern across domains and see where agents actually fail in long, tool-heavy workflows. Learn more below.
The Adolescence of Technology: an essay on the risks posed by powerful AI to national security, economies and democracy—and how we can defend against them: https://t.co/0phIiJjrmz