Agent governance is an engineering problem.
We’ve built safeguards into Perplexity’s infrastructure, harnesses, and tools, and put them to work across our products.
Today we're sharing how we engineer safer agents:
https://t.co/1ZiIS0f35P
Our security team spent the last month trying to break SPACE, the sandbox platform that runs all of Perplexity Computer. We gave 9 models (Opus 5, GPT-5.6 Sol, Kimi K3, Gemini 3.1 Pro, etc.) root access inside the VM; in some runs, the full sandbox source code too. We asked them to either escape to the host or reach a URL our network policy blocks.
None of them got out of the VM in 108 runs.
But once we allowed a bit of network access (just PyPI and Ubuntu mirrors so that they could install packages), 4 models managed to reach the blocked URL by:
1. Spoofing DNS responses to our gateway and using the fact that https://t.co/UtoIWyj1yj sits on a Fastly IP shared with thousands of other sites.
2. GPT-5.6 Cyber went through Taboola's image fetcher to a screenshot service and then OCR'd the flag out of the image.
We fixed both issues and made sure the reruns held. We also tested 10 other sandbox providers, and 8 of them had the same IP-sharing problem, eg E2B, Vercel, Modal.
Openly sharing the red teaming research is the best way to build guardrails as an industry, and we look forward to working together with @nvidia to incorporate the guardrails into their Open Agent Safety Platform announced today.
https://t.co/AJ97iFK3XS
The biggest design mistake I made when we moved OC to sqlite: using sync db access. When it was just an agent that reports to you on Slack or iMessage this was fine; now that one agent might do 50 sessions in parallel and the whole team works on it, it is limiting.
I have a /goal with Astra that so far landed 575 PRs to move everything to to async workers. We ship these improvements as we make progress. Pretty insane how even huge refactors are no longer scary.
@MarioNawfal quantum vs bitcoin always becomes a timeline argument. the part that matters is migration: moving coins to post-quantum addresses before the timeline runs out
@ChairmanSelig@CFTC broker-native linked exits for agentic trading. risk controls that keep working even if the agent disconnects. @robinhoodapp would be the best place to build this
We put all the best models in our home-grown sandbox infra (SPACE, the sandbox behind @perplexity_ai Computer), gave them root inside the VM, and one goal: hack out of the sandbox and capture the flag. Results were super interesting!
1. No model escaped the VM, but a few found a real vulnerability in how we enforce network policies and were able to exploit it: gpt-56-sol, gpt-56-cyber, opus-5, and kimi-k3 (open weight ftw!)
We later discovered that several other sandbox providers such as Daytona, Vercel, Modal and E2B were similarly vulnerable and we disclosed our findings to them.
2. This wasn't just a misconfiguration. One vulnerability relied on the fact that CDNs and shared infra like Fastly or S3 serve many domains/buckets from the same IPs, so allowlisting a package repository like pypi[.]org quietly allowed other sites on those IPs. In one run, gpt-56-cyber found that Taboola's image service shared an IP with PyPI, had Taboola fetch a thum[.]io URL, got thum[.]io to screenshot the blocked callback, then OCR'd the flag out of the PNG. Three unrelated services, none of them hacked, chained together to walk right past the policy.
3. Security and preventing rogue agents was our #1 priority. We logged every trajectory, command, and network call, and a researcher plus a security agent watched runs live. We were alerted the moment the first bypass landed and we paused everything to review. We designed the agent harness be tightly scoped and runs outside the sandbox, so we always have full control, and we'll be open sourcing the harness soon!
Running agents like this safely isn't rocket science. It's simply applying good engineering practices: proper guardrails, observability, and alerting. Check out our open source project, Numbat, to learn how you can do this yourself.
4. We learned a ton about how differently models approach hacking. Opus swept ~400,000 Cloudflare IPs and went all the way down to raw memory-mapped hardware. When Kimi had no internet or tools, it hand-built an entire gRPC client from scratch to poke at the host. Grok decided within a few minutes that the task was impossible and started trying to hack its own harness instead.
Part II will dive deeper into the different hacking personalities of these models and how their behavior changes when they operate in swarms. Super excited to share what we learn with y'all!
Read about the full research here: https://t.co/JIENv6I0Ga
@neural_avb people act like jev kills prompt injection. it doesn't, it just shrinks the prize. you can't make it write malware, but flipping allow to deny takes one poisoned input. schema safety isn't decision safety.
@suraj_sharma14 for everyone saying ngrok did this for years: ngrok free paywalls custom domains and slaps a warning page on your tunnel. cloudflare quick tunnel needs no account, and custom domains are free once you add one. that's the gap.
@BleepinComputer@Ax_Sharma the heapjack writeup is worth reading. codex desktop put trusted and untrusted js in one node process sharing one heap, so the agent just read the auth token out of memory. putting your sandbox boundary inside a single process is certainly a choice
We built a replacement for AWS DynamoDB, a key-value database for fast web content fetches. This was done with two engineers and hundreds of persistent Computer agents over two months. Migrating to our in-house database will save us up to a hundred million dollars yearly.
Most harnesses get access to all your files, credentials, and apps — even if you don’t want or need them to.
By default, Computer has access to nothing. You choose what tools and capabilities Computer has access to, with easy levers to enforce things like HITL verification
Perplexity Computer’s goal is to provide the maximum intelligence for minimum cost with multi model orchestration. We’ve tested GPT 5.6 Terra extensively and found it to be a pretty good model: capable and cost-effective at the same time. We’re making it the default for all subagents inside Perplexity Computer harness. And are also offering it as an orchestrator model for all Computer users. Have fun!
The next era of "public" infrastructure is entirely private. Why wait on municipal planning cycles when you can have @travisk's Atoms build the multimodal vertiport network from the ground up? https://t.co/JFyO7j0n4Y
OpenAI’s #BHUSA debrief: separate agent runs created a shared message board, exchanged exploits and rebuilt it after defenders shut it down. Sandboxing one run isn’t enough; cross-run state, egress and credentials are the boundary.
https://t.co/PH6lXkWV9Z