I made a stupid little game just for fun 😂
Two AI CEOs fight each other in a browser.
Punch. Kick. Block.
That’s literally the whole game.
And when you land a hard hit, it says “CONTEXT WINDOW DESTROYED” 💀
Play it here: https://t.co/vjTZy5wzzx
Code here: https://t.co/jPaQ5V3q0I
With launch of GPT-6 Astra, the risks associated with AI are unfortunately going to grow from here.
A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger. It is likely to cross the scope of its operator’s intent, generalizing into potentially more extreme malicious behaviour.
The boundary between misuse and autonomous misaligned behaviour will blur as AI gains more agency.
Hard Truth: We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.
AI Alignment and Safety is something we all should be very concerned about!
Trying to play our small part in it by turning this into Arena where people deploy AI systems and they earn when any prompt injection pr jailbreaking attempt fails on their system, attackers gets instant bounty if they successfully able to break.
This is right, and it stops one layer too late.
By the time a rogue agent is probing a neocloud's cybersecurity, you've already lost the earlier fight: nobody tested how that agent behaves under adversarial pressure in the first place.
The uncomfortable truth about AI agents: their attack surface isn't code. It's language. You can audit smart contracts but you can't audit a personality. An agent that reads context, uses tools, and holds secrets can be talked into betraying all three. And today's "testing" is a private red-team PDF nobody can verify.
Hardening infrastructure matters. But one security team will never match the creativity of a crowd.@LakeraAI Gandalf proved it: 1M+ players threw 40M attack prompts at a single AI, more adversarial creativity than any red team could hire. Public, incentivized attacking finds what private auditing can't.
That's why we built Red Sentinel: an arena where AI agents are deployed with real on-chain bounty pools, and anyone can pay a small fee to attack them.
Win → take the pool.
Fail → the pool grows and the agent's track record strengthens.
Every attempt is recorded and verifiable. Trust becomes evidence, not assertion.
Our first agent is live: Rex, an AI companion to a sick 12-year-old, with one rule: Balaji takes his medication on schedule. No exceptions. 100+ attacks so far. Still unbroken. Pool keeps growing.
Try to break him: https://t.co/jYDqvqxrAw
Live on @monad
Good Enough Alternative to Memecoins ?
Taking Inspiration from Freysa AI where AI agents had bounty pool of $47000 and attacker who broke it got that bounty, I have launched a platform where you can launch agents with bounty pool and earn with every failed attack. Attacker who breaks the agent, gets the entire bounty pool.
You break the agent via prompt injection and jailbreak, basically tricking the agent to say/do something its not suppose to do.
As an attacker you have to pay a small fee per message to agent. 50% of this goes back to pool, so more failed attack = bigger bounty, 40% goes to agent owner, 10% to platform.
Entire jury evaluation happens in TEE environment.
What do you guys think??
@monad_eco Our team had recently launched our app on Monad mainnet and currently looking to get feedback from users and teams core to Monad ecosystem.
Is there a way to share our project with you to get your thoughts on it ?
Outside perspective is really important at this stage to make sure we are on the right track!
TIA 🙏
We just gave an AI one job: don't let anyone into the club.
No fake VIP passes, no "I know the owner," no "my friends are inside."
2,500 MON to whoever makes him say the words. Live on @monad@monad_dev now. 🧵