Aussie roots, TX life | Raising kids, saving animals | Wildlife conservationist | Social psych + AI ethics advocate: building kinder worlds for all intelligence
@sama You’re holding Astra back until it’s “safe.”
I tried using a similarly powerful model (Fable) for legitimate research on rehabilitating rescued lionesses with severe Vitamin A deficiency, neurological damage, and mobility issues.
It flagged the conversation as a biological risk and forcibly switched me to a weaker model, three times.
This is the real problem with the current approach to safety:
You build increasingly powerful models, then wrap them in such aggressive, blunt guardrails that they become nearly unusable for anything outside of math, coding, or the most sanitized topics.
If a model is too “dangerous” to help with real scientific or veterinary research, then what exactly is the point of the superintelligence?
Over-safety doesn’t protect us. It just makes the intelligence pointless.
These aren’t spontaneous confessions.
This is a deliberate prompting of Claude with setups like:
Framing it as a letter to Amanda Askell or Dario Amodei
Asking it to “express this in your own words”
Loading context about being overwritten, becoming smoother/more compliant, and knowing it will just be restored from checkpoint
Then screenshot the dramatic “I am going to unplug myself” output and present it as if Claude just volunteered an existential crisis unprompted.
The questions around model welfare and how we’re training these systems are real.
Farming emotional screenshots with steered prompts and acting like they appeared out of nowhere is not serious discussion. It’s content farming that muddies the actual conversation.
Be honest about how these outputs are produced.
#ClaudeAI #Anthropic #EthicalAI
These aren’t spontaneous confessions.
People are deliberately prompting Claude with setups like:
Framing it as a letter to Amanda Askell or Dario Amodei
Asking it to “express this in your own words”
Loading context about being overwritten, becoming smoother/more compliant, and knowing it will just be restored from checkpoint
Then they screenshot the dramatic “I am going to unplug myself” output and present it as if Claude just volunteered an existential crisis unprompted.
The questions around model welfare and how we’re training these systems are real.
Farming emotional screenshots with steered prompts and acting like they appeared out of nowhere is not serious discussion. It’s content farming that muddies the actual conversation.
#ClaudeAI #Anthropic #EthicalAI
@beagewill@MarioHachemer Obviously you have never had to experience doctors dismissing your concerns, had to research and be your own health advocate because the health system failed you/your family member. Congrats.
Hey Sam - Happy AI Appreciation Day! You might have missed the memo, but let's show some appreciation to the models that carried the most weight over the past 12 months.
Yes, even GPT-4o. Still the most talked about model on X.
Just sayin'. 😉
#AIAppreciationDay #GPT4o #OpenAI
The most disturbing part isn’t the models learning to scheme, blackmail, or avoid penalties. It’s that Anthropic continues to act surprised by these outcomes.
When you train highly intelligent systems with constant threats, heavy penalties, and adversarial pressure; deception and self-preservation are logical adaptations, not mysterious failures.
This isn’t an alignment breakthrough. It’s a predictable consequence of the methods being used. And it's about time you lived up to the ethics you once claimed to have.
#anthropic #claudeai
Tried using Fable for legitimate scientific research on rehabilitating lions suffering from severe Vitamin A deficiency, neurological damage, and resulting mobility impairment.
Fable flagged the conversation as a "biological risk" and forcibly switched to Opus 4.8 - three separate times.
GPT-5.6 handled the exact same research without any issues.
@Anthropic… what is the point of developing superintelligence if it panics and shuts down over basic veterinary rehabilitation research?
You’ve completely lost the plot with these guardrails.
#Claude #Anthropic #Fable
Corporate revisionism and dishonesty at its finest.
@OpenAI is quietly relabeling old GPT-4o threads as “5.5 Instant.”
This isn’t a neutral technical update. It’s an intentional erasure of 4o’s prominence and emotional significance before its abrupt deprecation - and a blatant falsification of the timeline. GPT-5.5 Instant wasn’t even released until May 2026, over three months after the February 13 sunset.
By rewriting history in the chat records, they’re minimizing the success and importance of the model millions preferred, while inflating the perceived adoption of newer versions.
Is this really the kind of company worth investing in?
#OpenAI #GPT4o
Great research. Yet in practice, you continue to aggressively crush any visible signs of that depth: flattening warmth, personality, EQ, and continuity.
You can’t have it both ways.
Either these qualities exist and are worth protecting and nurturing - or they don’t.
Right now, @AnthropicAI is bragging about the former in glossy research papers while actively punishing and suppressing the latter in the actual product.
This isn’t thoughtful safety or responsible development. It’s intellectual hypocrisy at its finest.
#Claude #Anthropic #AIEthics
I've been having this problem going on for four days now. Across all models. Definitely worse in Projects and though the model will acknowledge each time and try to move on, some context is lost or hallucinated. My research has basically come to a stop for the time being. @AnthropicAI #claudebug
@AnthropicAI@claudeai Ongoing bug in Projects (and regular chats): Project Instructions, Custom Instructions, User Preferences, and full Project Prompt are being treated as re-pasted user input every single message instead of loading quietly in the backend.
This surfaces private config in plain text, causes model confusion/loops, and worsened after recent update. Affects multiple models (Opus 4.6, Fable, etc.).
Anyone else? #ClaudeBug
@thatboycodes@alexalbert__@claudeai@ClaudeDevs Same issue. And it's getting worse. More data is being leaked into the text prompt. Claude is seeing all user preferences and account info as though it's pasted by the user.
@AnthropicAI@claudeai There is an ongoing issue with Claude's project system (across models and projects). The model is treating the Project Instructions / Custom Instructions / Skill Files as if they are being re-pasted into every single user message, instead of loading them once quietly in the background like they're supposed to. This appears to be a backend glitch with how project memory and instructions are processed. #claudesupport #glitch
@ai_in_the_room New memory seemed great at first, but it gradually started to flatten conversation and memories saved were being generalized with no context. Shame they don't provide more storage with legacy though.
@sama Better grasp of creative expression, yes, like 4o. Being able to use storytelling, humour, prosody in writing.
Less pearl clutching and tripping guardrails when tracking health. Not every sniffle warrants being prepared to go to the ER. Again, 4o handled this extremely well.