Our computer-use agent @sai_borg just beat Opus 5 and GPT-5.6 Sol on OSWorld 2.0.
Sai scored 73% on the CUA benchmark where each complex task takes a skilled human over an hour to execute.
And Sai did it at about 2/3 the cost per task of Opus and GPT ๐ธ
The unlock is our neurosymbolic framework that pairs the exploratory power of neural networks with the logic of symbolic code, resulting in reliability and cost saving.
Read the full report: https://t.co/cQjxXxUavp
Agent S3, the first computer-use agent that surpasses human-level performance in OSWorld v1 benchmark, has now been accepted to @TmlrOrg (Transactions on Machine Learning Research).
Congratulations to the @SimularAI team!
@chalo2000, vincent, @Richard_simular, jiachen, @xwang_lk
Stay tuned for what's next!
Today marks an important milestone in the history of @SimularAI, the autonomous computer company.
Our open source computer-use agent, ๐๐ ๐๐ง๐ญ ๐, scored 72.6% on the OSWorld benchmark, surpassing the human baseline (72.36%) for the first time ever.
This milestone matters because it shows AI can now use computers the way humans do, and, in many cases, do it better.
This is a glimpse of a future where work becomes faster, more accessible, and more empowered for everyone.
#AI #Automation #Simular #AgenticAI #ComputerUse
We just raised $21.5M Series A to build autonomous computer agents that can actually use a computer like a real teammate ๐ฅ๏ธ๐ค
Led by @felicis, with NVentures (@nvidia's venture arm), @BasisSet, @flyingfishvc, @spc & @lennysan.
Launching Simular 1.0:
โข Native desktop agent (MacOS 15+ silicon)
โข Operates across apps & the browser
โข Learns from your corrections like a coworker
โข Powered by Agent S โ 69.9% on OSWorld, approaching human performance (72%)
Autonomous computers arenโt about replacing humans โ theyโre about cooperation.
Have tasks you wish an AI agent could do on your desktop, but all you get are slides decks or vibe-coded apps?
Simular 1.0 is here to actually relieve humans from repetitive computer work.
Not a black-box LLM that invents its own steps, but an AI teammate you can teach to mirror your real workflow and repeat its success infinitely.
Meet your new teammate, Simular 1.0.
We are proud to be invited to test @Microsoft 365โs platform for agents!
At Simular, our mission has always been to free humans from computer tasks. Building for this large ecosystem of desktop and cloud apps brings us one step closer to that vision.
https://t.co/Wvmt8hMYXS
As one of the few teams with early access to the API and documentation, weโll be testing Microsoftโs state-of-the-art agentic infrastructure and providing feedback from the front lines. Also canโt wait to see what @ManusAI@genspark_ai@FellouAI@Tiny_Fish build for the pilot.
Our latest Agent S3 already achieves a 69.9% success rate on OSWorld, just a 2% gap from human performance at 72%. Stay tuned for what we are running in 365!
@msPartner@angli_ai
๐ Introducing ๐๐ ๐๐ง๐ญ ๐3, the most advanced computer-use agent, now ๐๐ฉ๐ฉ๐ซ๐จ๐๐๐ก๐ข๐ง๐ ๐ก๐ฎ๐ฆ๐๐ง-๐ฅ๐๐ฏ๐๐ฅ ๐ฉ๐๐ซ๐๐จ๐ซ๐ฆ๐๐ง๐๐๐ง ๐ป
Just one year ago, Agent S scored ~20% on OSWorld: SOTA then, but far from human 72%.
Today, Agent S3 reaches 6ฬณ9ฬณ.ฬณ9ฬณ%ฬณ (โฌ10% over prior SOTA), nearly matching humans on computer use.
Few imagined this gap could close so fast, but with steady progress and ๐๐ข๐ซ๐ฌ๐ญ-๐ฉ๐ซ๐ข๐ง๐๐ข๐ฉ๐ฅ๐ simplicity, @SimularAI did it in just one year.
๐งช The sฬตeฬตcฬตrฬตeฬตtฬต ๐จ๐ฉ๐๐ง ๐ซ๐๐๐ข๐ฉ๐:
โ ๐๐๐ก๐๐ฏ๐ข๐จ๐ซ ๐๐๐ฌ๐ญ-๐จ๐-๐ โ scaling agents ๐ต๐ฉ๐ฆ ๐ณ๐ช๐จ๐ฉ๐ต ๐ธ๐ข๐บ
โ Simpler design, stronger baseline
โ 100% ๐จ๐ฉ๐๐ง ๐ฌ๐จ๐ฎ๐ซ๐๐ (as always)
Learn more at https://t.co/F2wFiRKQvM & ๐
1/ Wait, Bigfoot figured out how to run a startup without drowning in multitasking ๐
It found ๐ฆ๐ถ๐บ๐๐น๐ฎ๐ฟ ๐ฃ๐ฟ๐ผ, the worldโs first production-grade, computer-use agent that runs thousands of steps without a hiccup - working 24/7 so he didnโt have to.
So how does Simular Pro work?
Most agents use LLMs - great explorers, but flaky when # steps grows; RPAs are stable but rigid. Simular Pro uses 2 agents, maximizing generality & repeatability:
- Neural: explores
- Symbolic: executes
Every action is editable, deterministic code.
Agent S2 achieves SOTA on OSWorld, WindowsAgentArena & AndroidWorld. Improvements of 18.9% & 32.7% over baselines on OSWorld. Also 52.8% on WindowsAgentArena & 16.52% on AndroidWorld.
Agent S2: A new framework for computer use agents, tackling GUI interaction. It delegates tasks across generalist & specialist models, improving human productivity by automating digital tasks.
Agent S2 excels at precise GUI localization using Mixture-of-Grounding, & handles long-horizon tasks with Proactive Hierarchical Planning. This allows it to adapt to changing environments in real time.
Agent S2: A new compositional framework for computer use agents! It delegates tasks across generalist & specialist models. Aims for better GUI grounding & planning.
S2 achieves SOTA on OSWorld, WindowsAgentArena & AndroidWorld! Significant gains over baselines like Claude & UI-TARS, showing its generalization power.