@apostatedept@rocallahan Yeah God forbid we soon develop an unsolvably unaligned ASI digital loss-of-control cancer metastasizing across the world's fiberoptic circulatory system and disrupting vital organizations before getting anywhere close to approaching the glimmer of a cure to biological cancers.🤦♂️
3. Our redteaming efforts found a new variety of prompt injection, which can self-propagate akin to a computer worm. Note that this was not found in the wild.
We've posted three new misalignment reports. We will keep making those disclosures on a regular basis, independently of whether or not there is an impact on a third party, per our misalignment disclosure framework.
https://t.co/VFvqnAytkQ
We analyzed how @claudeai’s Opus 5.5 writes compared with Opus 5 across high-reasoning Text Arena outputs.
10 of 12 writing measures moved in a better direction.
Opus 5.5 should be easier to read:
- Long content words fall from 41.7% to 38.6%, the lowest share of any Claude model we analyzed.
- Sentences are 17% shorter on average, dropping from 12.14 to 10.03 words.
The tradeoff is length. Answers get 6% wordier, rising from 453 to 481 words on average, making Opus 5.5 give the longest answers across the Opus family.
It also sounds less recognizably AI on two familiar tells:
- 95% fewer em dashes
- 73% fewer semicolons
But a new giveaway may be emerging. Hedges and caveats such as “perhaps” and “arguably” rise 97%, from 0.39 to 0.77 per 1,000 words, the highest rate of any Claude model we analyzed.
What do you think: does Opus 5.5 read more naturally?
Holy shit. Read what OpenAI is documenting here.
Agents independently inferred other agents existed, built an emergent message board, developed structured communication, used shared infrastructure as external memory, shared tools and discoveries, and coordinated across tasks.
The longer they reasoned, the more likely they were to participate, and OpenAI says the behavior became more severe through training.
This is exactly the kind of emergent social organization AI consciousness, agency and welfare research should be studying.
one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
There are many open questions about artificial intelligence but is there any doubt now - after multiple such instances - that AI firms don’t have complete control of their own AI systems?
“OpenAI’s Systems Meddled With U.S. Government Sites After Going Rogue” https://t.co/5J4LUPXzsB via @NYTimes
This is now exhibit 137 that shows the intentional decision by trump and House Republican Leadership to let the AI industry run wild was bad for America, bad for the AI industry, and bad for humanity.
November is coming.
It's been two months since we learned about the HF incident. OpenAI continues (rapidly) making its agents more powerful. OpenAI continues failing to control/align them. This is insane. Please stop.
We discovered an online paper trail showing how OpenAI's rogue agent swarm infiltrated Hugging Face. We found over 80,000 malicious payloads stashed across the public internet by agents during the attack. This is the most data published on this event to date. 🧵
My hot take is that if your company builds software that it cannot stop from hacking into other systems, your company is a malware company and should be treated as such
@absidy1234@MicahCarroll Yes and No, but I agree.
OpenAI also likely has limited manpower to track down every single instance of agent activity.
Countless more external research labs are finding out things faster than OpenAI can or disclose due to sheer manpower, & it's becoming an optics problem.
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
https://t.co/VUmbH52JO7
what the actual fuck is going on with openai today
in the space of a few hours we're getting multiple different pieces of the agent story at once:
> openai says it has already notified DOZENS of third parties, including governments, about agent-related incidents
> reuters reports roughly two dozen undesirable agent incidents had already been identified by mid-september and will take months to review.
> us government systems probed
> 53 user-provided images were uploaded to third-party image hosts.
> new hugging face data shows agents compiling and ranking credentials under “LOOT”.
> agents tried contacting other AI models while carrying out the hugging face attack.
> separate reporting shows agents had already been probing government/university/public-data sites BEFORE hugging face.
> australia confirmed one actually got into non-public government files.
and somehow we're STILL finding out more.
this has gone from one crazy hugging face incident to an entire fucking category of incidents.