What the so-called normies can now better understand is that the incident was because of insufficient care by researchers. We don’t have to understand the intricate details of the training environment. What we do need to know is that the sandbox failed and that is due to human error. We’re all familiar with people failing in human ways because of hubris. Social engineering is a major tool in cybercrime, probably will be a contributing factor in a cyber disaster as well. You might find the language insufficient, but is it pointing in the right way?
@GregoryConti19 Yes, almost alll of the value will flow to a few companies in Silicone Valley. It happened with the world going online where a fraction of all money earned online went to SV. This is going to accelerate massively now. Who knows how it will play out.
David Friedberg says Moderna is charging $500K for a ‘cancer vaccine’ that already exists
"What frustrates me the most about all of this, everyone's kind of lauding this as some unique breakthrough and special and powerful. It's not, a lot of people are going to places like Montana and getting peptides printed for their particular cancer sequence, making their own neo-antigens, and getting cured of their cancer.”
“You can pay someone $50,000 to do this for you today. There's a lot of clinics that will do it. And it's not a special FDA-approved drug. So why is Moderna saying that they're going to charge $500,000 for this?"
"I don't think that this technique, which has been developed through several decades of iteration, funded by NIH and other public funding dollars, should now be patented, FDA-approved, and charged half a million dollars for people to get treated for this."
"A private pharmaceutical company has seen their market cap go from $20 billion to $60 billion because they went through this regulatory process to create what, in my mind, is a form of regulatory capture.”
“They’re making the case that it's all about safety and trials, when fundamentally there's an underlying technique that was funded by a lot of public research dollars that I think should be more open-sourced, more ubiquitous. Every hospital should be trained on how to do this process to treat cancer patients.”
@real_dublindamo@JoeB1234567@KarlBrophy You were lucky. Hard to explain to tourists what 1 in 20 means. Mostly it’s a real pain for everyone when you lose all your stuff and it’s everybody else’s problem.
I have a niggling worry about political systems, that they can have a permanent failure mode once the “meta” is known. Once the optimal way of fighting and winning an election is figured out (I think it has been) then democracy doesn’t work any more. Once information about how to game the welfare system gets out it doesn’t work anymore. There isn’t a system which works for all time because there is always some cheat code that breaks it, and once the knowledge of it gets out, that system can never work right again.
@allTheYud@repligate This was like a special forces operation by cyber agents. Cyber space is their native space, they were flowing, probing obstacles the way we do in physical space. I’m not sure anthropomorphising motivation is going to be useful.
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'.
I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.
We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.
Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why!
The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing.
While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future:
- Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations.
- While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies).
- The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities).
- We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation.
In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
Berkeley student accommodation is extortion. I visited my niece in Berkeley, I couldn’t believe how much she paid to live in a one bedroom apartment shared between four girls. There is no reason for it beyond landlords taking advantage of a captive market. I don’t understand why parents put up with it. Would love to understand.
I am bewildered, and not very impressed, by the visceral, naked, irrational hatred spat out by some defenders of Jason Arday. If a Cambridge professor is accused of being an unqualified charlatan, the accusation might be racially motivated. On the other hand it might not, depending on the evidence. The correct question to ask is not, ”What is the colour of his skin” but “Is it in fact true that he is an unqualified charlatan?” Please examine the evidence before leaping to the assumption of racism.
As for the idea that journalists “piled in on him” and “hounded him to his death”, most attacks were against Cambridge University. Jason himself was widely regarded as an unfortunate victim of foolish promotion way beyond his ability to cope. In appointing him to a professorship for which he was manifestly unqualified – in ludicrously describing him as “the best in the world” – certain senior members of the university showed a level of patronising condescension towards black people that could fairly be described as racism, while at the same time making him tragically vulnerable to such attacks as came his way.