After almost 2 months of cooking, we probably have one of the most complex AI security datasets ever made.
> In the 2nd campaign, 10,987 users sent 757,484 unique prompts
> Consumer-scale, not researcher generated
> 125+ countries, multiple languages
Remember, we have users that paid to send prompts. Inverted incentive. Quality is incredible.
Stay around. It begins.
During our adversarial campaign, 1,211 distinct users fired 1,547 prompts of that attack class at our defender.
26 used the exact same tactic that tricked Grok to drain Bankr.
0 broke it.
We already have the solution for this. @xai let’s talk.
Perplexity shipped an AI agent that lives on your Mac full-time. It works across local files, apps, and the web.
Agents aren’t coming. They’re moving in.
Who’s testing what they do when no one’s watching?
Roughly 96% of enterprises are already running AI agents. Only 21% have a governance model to match.
This is the most dangerous gap in tech right now.
Palo Alto Networks just showed how a red-team agent got a financial copilot to execute a $900 withdrawal. No exploit. No breach. Just clever reframing.
Agents don’t get hacked. They get persuaded.
OWASP’s Q2 2026 landscape report names the top threats: prompt injection, agent privilege escalation, data poisoning, hallucination drift.
These aren’t theoretical. They’re happening in production.
Last week's AI Red & Blue Team Summit proved the point. Day 1: exploit live LLM workflows. Day 2: build detection rules.
The threat is real. The industry is waking up to it.
That's why the participation layer exists. Humans adversarially testing AI before it ships, not once, not on a schedule, but live and continuous.
Red-teaming shouldn’t be a line item. It should be infrastructure.
Claude Opus 4.7 now argues with you when you’re wrong and runs complex projects for hours unsupervised.
AI isn’t just getting smarter. It’s getting opinionated.
The better these models get, the harder they are to evaluate. Unless humans are in the loop.
AI is the most funded industry of our time, accounting for 50% of all global venture funding. It's also the least tested at scale. The gap is growing faster than the models themselves.
The next decade depends on closing that gap.
AI moves fast. Security must follow.
The gap between early and late is measured in shards.
Every prompt you send to King Arthur, every level you break, your stack grows. The leaderboard already knows who's been paying attention.
The best time to stack shards is now 🌀
The King is coming back.
Free to play for everyone.
Learn the model. Probe the edges. Level up.
Test Stormrae and get ready for the next releases.
Free prompts, more shards.