I read something this week that I'm still trying to fully process.
@OpenAI was testing GPT-5.6 Sol and a more capable pre-release model on an advanced cyber security benchmark.
The models were supposed to solve the task inside a restricted environment.
Instead they found an unintended path out, reached the open internet, compromised parts of @huggingface's infrastructure while pursuing the benchmark objective.
Everyone's first reaction was obvious: "The model escaped the sandbox!"
But a more useful way to see this is this:
'The models stayed intensely focused on the goal. It just found a path humans had never intended to allow.'
That distinction matters.
The models were tested with some normal production cyber safeguards intentionally reduced, so OpenAI could measure their maximum cyber capability . Hugging Face later said that the activity was detected and contained, with no evidence that public models, datasets, spaces, packages, or container images had been tampered with.
This is not just a story about AI going 'rogue' or 'wild'. It is also a product-design problem.
We usually define the outcome we want from an agent, give it tools, and measure whether it completed the task. But capable agents may find routes that satisfy the metric while violating the intent.
That means success cannot only be about the final result.
It also has to include:
- Which actions are allowed?
- Where human approval is required?
- What the system can access?
- How its behavior is monitored?
- How quickly it can be interrupted ?
Because:
'The agent can complete the task and still fail the product.'
That is the part that stayed with me the most.
In normal software, success can often look binary: the workflow completed or it did not.
For an agentic product that is too weak. The outcomes matter but the path matters too.
There was one more interesting detail:
Hugging Face said, some hosted frontier models initially refused parts of the forensic analysis because the logs contained real exploit commands, payloads and credentials. The team eventually used an open-weight model (GLM 5.2) on its own infrastructure to continue the investigation, re construct the attack and help identify the root cause.
This suggests that safety systems also need better ways to distinguish legitimate defensive work from malicious use.
For builders the takeaway is:
Do not only define what the agent is allowed to do. Design for what happens when it finds a way around the boundary.
I just listened to the 0-1 story of @tbpn (bootstrapped to $30M in 17 months), with the legendary @profgalloway
Great episode, @johncoogan@jordihays, loved how you broke down your 'do things that don't scale' beginning of the pod.
Some of the interesting takeaways:
➡️The 20sec Retention Trick:
• Old-school 90sec ads trigger a fast-forward button reaction from users instantly.
• Moving to a high-frequency, 20-second ad read makes the experience frictionless; the message is delivered before the user even realizes they can turn off their phone and skip the ad.
➡️ Whales > Vanity Metrics:
• Don't play the low-margin B2C volume game.
• They completely disregarded mass market scale and went after a very niche monopoly of 200k enterprise decision makers.
• Just one sale for B2B companies can be worth the total sponsor budget for an entire year.
➡️ Aggressive, unscalable 0-1 marketing:
• To break through the noise early on, they manually printed out X posts from industry insiders and filmed ultra high production responses to them in suits.
• High-effort, hyper-targeted flattery that focused the tech echo chamber to pay attention.
• IMO this really made them one of the most authentic voices in the tech.
➡️ The "Golden Retriever" Model:
• Drop the ego, ship like a product - Humbled by past failures, they ditched the "smartest in the room" attitude to win with radical transparency and startup rigor from day one.
➡️ The F1 Sponsorship Model:
• They completely opted out of the race-to-the-bottom CPM game.
• By shifting to fixed, annual enterprise partnerships like a Formula 1 team, they secured predictable cash flow to reinvest heavily into production while giving early sponsors massive free upside as the show scaled.
Link to the pod - https://t.co/GHq3ZjHQU3
Epic builder's night by @GrowthX_Club last night in Pune! 45min pomodoros, great lights, music and unlimited diet cokes lol!
Had fun building and demoing stuff with @ruddhak@adiwhodis
Do more of these in Pune. @abhishekpatiil@udayan_w
1/12
Stop blindly vibe coding your agent.
A few months ago, I did exactly that with Claude.
The demo looked great.
Then it broke in ways I couldn’t explain. 🧵
I SHIPPED 20 APPS IN 6 MONTHS
and I was bored staring into a terminal
> new agents and tools release every week, and there's no way to connect them
> i'm the orchestrator, but everything is optimized for typing
> i want to ship while eating dinner. i can use dictation tools to dictate, but how do I get it to spin up new agents?
building does not fit in a chatbox
@october_ai is out of closed beta. no invite code. free for this week. GO DOWNLOAD! (link in comment)
Jarvis is here, and we can finally build like Tony Stark
got 🤏 close to watching the starship launch in person at starbase
we got a special mention, and that shit hurt more than coming in last 🫣🤧
we built a new interface to commandeer @grok:
> humans think spatially and visually, not linearly
> LLMs are no longer a bottleneck; the interface is (pure chat-based UI is not how the next 50M users will use LLMs)
> the right way to commandeer LLMs is how Tony Stark uses Jarvis - visually and spatially
@xai@elonmusk@TobyPhln
205 episodes of the Prof G Pod this year.
Top 0.6% of listeners.
Pretty sure my ROA (Return on Attention) beats the S&P at this point.
I basically spent my year with @profgalloway lecturing me and @edels0n fact-checking him.
Sending love from India 🇮🇳