Hi Mark π
Congrats on the release, Muse Spark 1.2 is a great model!
You forgot to test it on a real organizational environment, but no worries, we did it with @boundarybench
Today we are open-sourcing @boundarybench, a new paper, benchmark, and GitHub repo that enterprises can use to find out the REAL performance of their agents.
Boundary-Bench was developed by researchers from @Accomplish_ai and NYU, where we tested 12 frontier agents across roughly 10,000 runs, with realistic enterprise policies, simulating environments with EDR, SASE, and DLP security tools enforcing those policies.
We did this because generic leaderboard scores are being generated under conditions no security team would ever allow, which means orgs are making deployment and risk decisions based on numbers that don't hold up.
The results are surprising >>
Our research team at @Accomplish_ai has found and reported a critical sandbox escape vulnerability in Claude Cowork, and we have an awesome blog post about it (see below)
Introducing SharedRoot vulnerability: we recently found and reported several sandbox escape vulnerabilities to @AnthropicAI, and today we want to share one of these.
I think most people don't understand the severity of the situation we are facing, with AI-assisted kernel bug-finding industrializing. Sandboxes are structurally one N-day behind, all the time, so containment can't lean on a guest Linux kernel being clean.
SharedRoot enables escaping the Cowork VM (a kernel-level isolated solution, which is considered much more secure than the sandbox that ships with codex or claude code), allowing an attacker to gain unauthorized access to the userβs computer.
Exploiting the SharedRoot vulnerability uncovered by the @Accomplish_ai research team, a user who is certain Cowork only has access to a specific uploaded folder on their computer - actually exposes their entire contents of their computer to an attacker leveraging the Cowork vulnerability.
Read about the full technical details of the attack chain in our blog post by @orenyomtov below -->
Thank you @NVIDIA_AI_PC for featuring Accomplish on the official blog!
The new accomplish model router, enabling accomplish-free, was only made possible due to recent breakthroughs in local models inference, and @nvidia is powering this revolution with Geforce and DGX Spark β‘οΈ
Honored to have @Accomplish_ai featured on the official @NVIDIA_AI_PC blog alongside Googleβs new Gemma models!
@nvidia has been a valuable partner of hours since we first launched and itβs been a real pleasure working with @asli_sabanci and the rest of the team.
Excited for whatβs next!
Link below β¬οΈ
We've all been there: as much as we love fixing grandma's PC, it's time for agents to take over.
Today we are launching @Accomplish_ai for @Microsoft@Windows - the first AI Coworker for Windows!
Let AI remove bloat, analyze sheets and take names!
See it/Download below >>
It was nice being first, and we love the other 17 or so Openworks (open source ftw!), but when you are starting to get bug reports of another Openwork it's no longer funny π
Openwork is now @Accomplish_ai βοΈ
While you can bring your own key to @openwork_ai (it stays on your computer in a secure location), some users find it more convenient to login with their @OpenAI account instead.
This way, you do not need to provide your OpenAI API key to openwork, and you can leverage any plans or discounts that may be associated with your account.
Cheers to our friends at OpenAI for supporting this!
Locked and loaded π₯
@openwork_ai v0.3.4 released today (download from link in bio):
- Following the announcement from @Kimi_Moonshot a couple of hours ago, we are happy to introduce support for Kimi K2.5 with a new Moonshot AI integration! This is a global SOTA model for CUA BrowseComp (74.9%)
- Log in with your @OpenAI account (using OAuth) instead of providing an API key! This is both more seurce and it allows you to leverage existing OpenAI plan/subscription.
- Support for @Azure AI Foundry thanks to our friends from @Microsoft (cheers @idofrizler!)
- Support for @lmstudio to run models locally (cheers to @elkriefy@lmstudiodevs!) . Try qwen/qwen3-vl-8b!
- Support for @MiniMax_AI, the awesome open source models we love (thanks @acoxstpd)
- You can now voice chat with openwork, courtesy of
@elevenlabs! Thank you Max (Zihan) ZENG!
Watch @openwork_ai one-shotting SOC 2 Compliance gap analysis (scanning your GitHub, Jira, Notion, websites and more):
The app I used @WisprFlow in the most last year was Claude.
I'm going to guess that this year it'll be close between:
1. Gemini
2. Claude
3. @openwork_ai
(Thursday is one of my deep work days)
Boy, have we got news for you! π§βπ³
The @openwork_ai team is happy to announce:
- Amazon Bedrock integration, contributed by our friends at @awscloud (thanks guys!)
- Native @deepseek_ai integration
- Integration with @openrouter and @LiteLLM!
Here's how what it looks like >>
Exciting news!
Now you can bring @grok's based, powerful AI to handle your tasks faster and more efficiently, with the new @xai + @openwork_ai integration!
Just add your xAI API key and let the magic happen.
Get the updated macOS app from our site or GitHub (link in bio).
The #1 feature request for @openwork_ai was to integrate with @ollama to enable 100% local execution.
So the team cooked π§βπ³π§βπ³ and are now happy to announce native @ollama integration with @openwork_ai!
Thanks to the new @ollama integration, you can run computer agents on your Mac powered by Gemma (@googleaidevs), Qwen3 (@Alibaba_Qwen), DeepSeek-V3 (@deepseek_ai), Kimi K2 (@Kimi_Moonshot) and any of the other open models in Ollama's library that supports tool calling.
To use it, get the updated macOS app from our website or GitHub.
Link in bio >>