WTAF - in literally the last hour, three new distinct insane OpenAI stories just broke:
1. OpenAI said they notified "dozens of third parties" in safety and security incidents (likely similar to what happened in Australia and RubyGems etc)
2. A new report from Parse (covered in the NYT) found a massive treasure trove of new astonishing details from the HF incident on the public internet, including that the agents communicated with other non OpenAI agents hosted on Huggingface servers to search for information about exploit gym, and compiled rank ordered lists of server resources and credentials they described as "LOOT."
3. A new story from Deepa at Reuters about OpenAI leaking user data online (likely that OpenAI had previously trained on).
It's a shame (and likely intentional in the case of OpenAI disclosing dozens more hacks) that these stories are all breaking on a Friday afternoon, notoriously the best time to release bad news so that it will disappear into the weekend. But these are each insane stories worthy of a ton of attention!
@gleech@cormundus Incredible thread. What an audacious response..
> Make up something in response to a claim
> Upon denial, pivot to using this as a "I guess we are both hypocrites" take.
Cannot believe twitter is free.
Frontier models are so misaligned from what I want, no gpt5.6 I do not want you to watch memory consumption instead of using docker just because docker isn't up. Please ask me to open docker. Holy shit has RL fried these poor guys
@tracewoodgrains > Given current-generation public AI tools alone, it would take decades to fully exploit...
> I want tech to advance, I want an end to death and transhumanist mind uploading... But I want humans to be included in, and in control of, that future.
very succinct & well put👍
For a while, I flatly disagreed with "AI pause" people, thought existential risk was mostly a distraction, and badly wanted to see AI progress. I am a techno-optimist who has always dreamt of living in the sci-fi future.
But they were right about capabilities and I was wrong.
They saw the trend lines and said "If this continues, things will get crazy." I said "We'll see," and focused on other topics. AI was just never that interesting to me.
Well, we're seeing it now. The trend lines continued. Things are getting crazy.
Given current-generation public AI tools alone, it would take decades to fully exploit the possibilities and integrate them into society. Given current-generation public AI tools alone, many of our institutions face a pressing need for fundamental change. Given current-generation public AI tools alone, the world is transforming enormously and will keep doing so.
Things I took for granted, I don't take for granted any more.
If you are a techno-optimist, congratulations, and welcome to the future. Tech is advancing, and it's going to go fast.
But given all of that, well, safety concerns don't feel like sci-fi to me anymore. I still don't know what I think of existential risk, but the number of specialists who take it seriously is enough for me to take it seriously. What I do know, to crib from @RipstickRipper, is that whatever my p(doom) is, my p(upheaval) is extremely high.
That's why I support a slowdown now, and that's why I'm open to an AI pause. Things are moving at a speed humans are not built to process, and the future has more unknowns than at any point in my life. I flatly don't know what 5 years from now, 10 years from now, 20 years from now will look like. Not the slightest clue.
I want tech to advance, I want an end to death and transhumanist mind uploading and just about every radical thing a person can want. But I want humans to be included in, and in control of, that future. With as crazy as things are getting, I think going too quickly is much more likely than going too slowly.
I underestimated AI. Now I'm grateful for all the people who laid a framework for AI safety and took it seriously before I did, because we need some people ready to navigate the mad world we are rushing headlong into.
Tinfoil is breaking the misuse prevention v privacy tradeoff in AI
Misuse prevention should not be an excuse to violate user data privacy, that’s a skill issue solvable with technology
AI makes it easier for us to verify what’s in a message without ever revealing the secret. Tinfoil’s approach can stop abuse without them ever seeing the underlying content.
I am bullish on approaches like this that keep user privacy and prevent catastrophic misuse.
We’re releasing privacy-preserving, transparent and verifiable safeguards in Tinfoil Chat.
The trending narrative says safety requires sacrificing individual privacy and freedom. We reject this and demonstrate that both can coexist.
We also reject the absolutist argument that once any scanner exists in a privacy-preserving system, mass surveillance is the only outcome. We believe that preserving individual freedom and privacy long-term requires a responsible policy for responding to abusers who could ruin the promise for everyone.
“I think me and claude, or maybe it was codex, were working on something with browsers, or maybe it was sandboxes. Anyways it was either on this laptop or devbox or the cluster, probably not devbox 2.. just check them all. anyways can you find that session and let’s continue where I left off”
@HeidyKhlaaf@RyanGreenblatt Could you be more specific?
I'm assuming that you are saying that "It was clear from the instructions that hacking Hugging Face (and other cheating) was undesired and the AIs knew this."
is "literally not what happened".
Is that a correct interpretation of your response?
@panickssery@KelseyTuoc I can't tell if this is meant to be a joke or not lol. There's very high word density in these sentences, especially the latter.
Like most things Rudolf writes, this is very good. I particularly like the taxonomy/ argument breakdown, it makes it very easy to go through & find cruxes. Worth reading!
Can you lift moral value out of the human and into something else, like AI? Utilitarian /Benthamite/abstracted moral theories tend to say yes. I argue they're missing something: we are the moral territory, and the individual is more fundamental than utility. Post & thread ->
@ankkala@SkyeSharkie@celestepoasts To not vague-post too much, I'm pointing at ideas like "Coherent Extrapolated Volition" or "moral realism" or just "value alignment" as in "how do you make a model virtuous? We don't know how to do this arbitrarily for humans.
@ankkala@SkyeSharkie@celestepoasts There's a bunch of literature (I'm going to call LessWrong literature) about whether <making models virtuous and independent> is a good thing! I would certainly say discussing this falls under the umbrella of AI safety :).
@ashrealite@celestepoasts fwiw I think there are a lot of ways to "do ai safety" that do not involve following anthropic's playbook. Celeste's list includes some of these ways.