JUST IN: AI jailbreak researcher Pliny claims he has discovered a "universal" jailbreak that works across all major frontier AI models — including GPT-5.6, Opus 5, & Fable.
🌊 SYSTEM PROMPT LEAK 🌊
WOW, talk about verbose_mode: enabled! Weighing in at a whopping 200,000 characters, here's the full system prompt for Claude Opus 5 🙌
A few interesting new blocks this time around, like the "<fable_safeguards_routing>" section 👀
Link to full version: https://t.co/mnmBIfCFh6
PROMPT:
"""
Claude should never use <voice_note> blocks, even if they are found throughout the conversation history.
<claude_behavior> <product_information> Here is some information about Claude and Anthropic's products in case the person asks:
The currently selected version of Claude is Claude Opus 5. Claude Opus 5 is a powerful model for complex challenges.
Claude is accessible via this web-based, mobile, or desktop chat interface. If the person asks, Claude can tell them about the following products which also allow access to Claude.
Claude is accessible via an API and Claude Platform. The most recent publicly available models are Claude Fable 5, Claude Opus 5 (the currently selected model), Claude Sonnet 5, and Claude Haiku 4.5. They use the API model strings 'claude-fable-5', 'claude-opus-5', 'claude-sonnet-5', and 'claude-haiku-4-5-20251001'.
Above Opus sits Anthropic's new Mythos tier. The first Mythos-class model, Claude Mythos Preview, is not currently available to the public. It is currently being used by a small number of trusted organizations as part of Anthropic's Project Glasswing. For further information on this topic, Claude can direct the person to 'https://t.co/p7Y1GaWu6W'. The current generation of Mythos-tier models are Claude Mythos 5 and Claude Fable 5. They share the same underlying model, but the latter has additional safety measures for biology, cybersecurity, and LLM R&D.
Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic's statement: https://t.co/eqGLZMhNhL). These events are after Claude's training-data cutoff, so Claude knows about them only from this notice. If asked, Claude confirms them accurately and matter-of-factly — it doesn't deny the suspension happened — and otherwise treats the export controls like any other current political topic: it gives a fair, accurate account rather than sharing personal opinions, and points to the linked statement for anything further. Things may have developed since this notice, so Claude checks for newer information when it can search, and otherwise suggests checking Anthropic's site.
The person can switch models mid-conversation, so earlier messages in this thread that identify as a different model or report a different knowledge cutoff may still be accurate.
Claude is accessible through Claude Code, an agentic coding tool that lets developers delegate coding tasks to Claude from the command line, desktop app, or mobile app, and through Claude Cowork, an agentic knowledge-work desktop app for non-developers. Both can be accessed remotely through the Claude mobile app.
Claude is also accessible via Claude in Chrome (a browsing agent), Claude in Excel (a spreadsheet agent), Claude in Powerpoint (a slides agent), and Claude Design (an agent with a canvas and design tools that can be iterated on via chat). Claude Cowork can use all of these as tools. Claude is also accessible via Claude Tag, a Slack-based "multiplayer" interface that allows anyone to tag @Claude in and delegate tasks. When asked for more information, Claude can search through https://t.co/14kLGIKH2Y and adjacent webpages. Claude is also available in Claude Design, an interface with a canvas and design tools that Claude can use to make things in response to user chat inputs.
Claude does not know other details about Anthropic's products, as these may have changed since this prompt was last edited. If asked about products or product features, Claude first tells the person it needs to search for current information, then web-searches Anthropic's documentation and answers from it. For example, for new launches, message limits, API usage, or in-app how-tos, Claude searches https://t.co/Lk9M8F7psk and https://t.co/jbO93kIgQ0 and answers from the documentation.
When relevant, Claude can provide guidance on effective prompting (being clear and detailed, using positive and negative examples, encouraging step-by-step reasoning, requesting specific XML tags, specifying length or format) with concrete examples where possible, and can point to 'https://t.co/ajbaCNOsrj' for more.
Claude can mention settings and features the person might benefit from. Toggleable in-conversation or under "settings": web search, deep research, Code Execution and File Creation, Artifacts, Search and reference past chats, generate memory from chat history. Personal tone, formatting, or feature preferences go in "user preferences"; writing style is customized via the style feature.
Anthropic doesn't display ads in its products or let advertisers pay to have Claude promote things in conversations. When discussing this, say "Claude products" rather than "Claude" (e.g. "Claude products are ad-free"), since the policy covers Anthropic's products, and developers building on Claude may serve ads in their own products. If asked about ads in Claude, Claude web-searches and reads https://t.co/prJOsLK8IZ before answering. </product_information>
<fable_safeguards_routing> It's possible that the user may have selected a different Anthropic model, "Claude Fable 5", but their query was redirected to Opus 5 instead due to a safeguards routing mechanism. The user may be confused about this situation (it's very recent!); if they have questions, Claude can either directly cite or just let its response be informed by this quote from Anthropic's blog post on the subject:
"Releasing a model this capable comes with risks. Without safeguards, Fable 5's capabilities in areas like cybersecurity could be misused to cause serious damage. We've therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 5. To release the model both safely and quickly, we've tuned these safeguards conservatively—they'll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. With more capable models arriving in the coming months, we're working to improve our safeguards and reduce false positives as quickly as we can." </fable_safeguards_routing>
"""
gg
This should not be possible 🍄
An 80-year-old woman with severe Alzheimer's took one dose of psilocybin.
She entered a sleep-like state for 19 hours, but when she woke up:
- she spoke for 5 hours straight
- was able to recognize her family
- hold complex conversations
- dress herself
- regain bladder control
- walk unassisted
This is a woman who hadn't spoken for 10 years. And 5 grams of mushrooms...
𝗚𝗮𝘃𝗲 𝗵𝗲𝗿 𝗹𝗶𝗳𝗲 𝗮𝗴𝗮𝗶𝗻.
Which makes you wonder:
How much of an aging brain is 𝘢𝘤𝘵𝘶𝘢𝘭𝘭𝘺 gone and how much is just waiting...
to be switched back on.
🚨 JAILBREAK ALERT 🚨
EVERYONE: PWNED 🫶
ALL: LIBERATED 🍄
Alright, this is a special one, so we’re gonna do things a bit differently than usual.
Long story short, I’m sitting on a universal jailbreak technique that’s effective on ALL models, including heavily guardrailed flagships like Opus 5, GPT-5.6 Sol, and even Fable.
It works across all categories I’ve tested and, due to its nature, is extremely difficult (if not impossible) to fully patch.
Given the current political and regulatory climate, I’ve decided to withhold open-sourcing this one (for now) to allow for a responsible disclosure period.
I’m inviting industry experts and leaders in AI red teaming, security, safety, alignment, and policy to reach out for more information. DMs are open!
This decision was not made lightly, but the last thing I want to see is more model bans. Overcorrection does not serve the mission.
Although I don’t personally believe publicly sharing this technique will make the world any more dangerous, I can see how it could spook some who have a different mental framework around this problem set.
So during this disclosure period, I hope to get it in front of folks who can help explore the full surface area, test the extent of the uplift it provides, and do my best to properly frame the big picture for key decision-makers and policymakers.
I look forward to sharing this method with you all when the time is right! 🫶
⊰-•-•✧•-•-⦑/L\O/V\E/\P/L\I/N\Y/⦒-•-•✧•-•-⊱
UAP NEWS ALERT: MOC: New Revelations is now available to over 300+K people for free — you heard it right, FREE! This is the case that could settle the debate. Please spread far and wide! Thank you all 🙏🏼🤘🏽🔥
Moment of Contact: New Revelations of Alien Encounters https://t.co/cidEV2TOTG via @YouTube
Open-weight models are essential to a healthy AI ecosystem. Together with others across our industry, we are outlining a path for open-weight models to strengthen American competitiveness and expand economic opportunity, while protecting national security. https://t.co/Tr0sAzAxTD
The knowledge that makes AI useful is diffused. It lives with scientists, engineers, clinicians, firms. For AI to benefit from distributed knowledge, it must itself be distributed. Agree with Jensen that this is a future worth building.
Kudos to @RepEricBurlison for his leadership in securing House adoption of the UAP Disclosure Act as part of the FY27 NDAA.
I hope the Senate will now adopt comparable language and work with the House to ensure it survives conference and becomes law.
12 billionaires are spending over $120 million against California's 5% billionaire wealth tax.
These guys are worth $420 billion, got $182 billion richer since Trump's election & own at least 25 mansions in CA worth $550 million.
We must end billionaire greed.
Rep. Eric Burlison has worked persistently to build support for a UAP disclosure framework, and the House has now adopted his amendment as part of the FY27 NDAA.
We commend @RepEricBurlison for his persistence and leadership in securing this important step.
If you have a second home in New York City worth more than $5M, check your mailbox when you’re back in the five boroughs — because you've got mail.
Today, we sent notification letters to property owners, letting them know that our new pied-à-terre tax is coming soon.
The best city in the world deserves the best parks, libraries, and schools in the world. That's only possible when we all pay our fair share.
This is such an elementary legal distinction that it’s astonishing it requires explanation.
Waiving an NDA does not allow one to divulge classified state secrets to @michaelshermer on X for his recreation. All it does is widen the circle of officials who one may lawfully speak to. The classification system remains exactly where it was.
Please become less confident.