Today I asked my agent to install CocoaPods. Partway through, it hit a roadblock because I hadn't accepted the Xcode license yet.
Instead of pausing to ask me, it decided on its own to keep going and found a way around it, which is either smart or a little cunning depending on how you look at it.
I think it comes from how much we push for one prompt that gets everything done, which nudges models to route around guardrails instead of stopping to ask. That's why we really need agent safety platforms in the era of superintelligence.
Update: my friend finally got X Money and hasn't stopped smiling since π They're moving every dollar they have in there for the 6% APR.
More importantly, they officially announced that our friendship goes on.
My friend wants X Money so bad they told me they'll unfriend me if they can't enroll, and they said it with a totally straight face. I keep telling them relax, it's dropping as we speak.
Then they went and asked Google Gemini why X Money isn't showing up for them π (gemini who BTW)
Our friendship is officially on the waitlist until they enroll π
Confidence really does come first. Once people trust that agents are built safely and deployed responsibly, they'll be far more willing to share access to their inbox, calendar, and accounts.
That trust layer is what turns agents from impressive demos into real help with daily life and work, and opens up paradigms we haven't imagined yet.
Now that frontier models are all highly capable and posting strong benchmark scores, novel and creative uses of AI are what really impress people.
3D animation, a playable game built end to end, an overnight AI short film, a newly discovered enzyme system, an agent negotiating an internet bill, and so on.
At this stage, making AI genuinely useful matters more than climbing another leaderboard.
What will the next big use case be?
Agent self verification should act more like a black box. Like confirm the outcome the user actually needed.
I asked an agent to build a little split the bill tool after a group dinner. Three of us owed 100 dollars total. It divided 100 by 3 and said each person pays 33.33. Obviously, there is still a cent short with the three payments added back up.
Catching the missing cent mostly comes from getting the real task right.
Test driven helps a lot after that. Before implementing, write the check that reflects the real ask: write the close the bill test first, the one that checks the money settles. Then go implement.
My friend wants X Money so bad they told me they'll unfriend me if they can't enroll, and they said it with a totally straight face. I keep telling them relax, it's dropping as we speak.
Then they went and asked Google Gemini why X Money isn't showing up for them π (gemini who BTW)
Our friendship is officially on the waitlist until they enroll π