Brecha entre capacidad y supervisión: la capacidad de sistemas de IA para perseguir objetivos complejos y potencialmente desalineados aumentan más rápido que nuestra capacidad —incluso ayudándonos de otras IAs— para entender y vigilar su comportamiento.
https://t.co/yFlssMcTCB
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'.
I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.
We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.
Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why!
The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing.
While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future:
- Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations.
- While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies).
- The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities).
- We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation.
In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
La muy incómoda sensación de sentir que el mantra defendido durante estos cuatro años era "primero desarrollemos todo lo que podamos y luego ya veremos" a pasar ahora a un "¿pero por qué se ha estado desarrollando todo sin ver primero qué y cómo?"
this is a critically important moment for cyber defense with AI; there is not much time to act.
we are happy if you want to work with us or any of our competitors or partners, but please take this moment seriously.
only an urgent and intense collective response will work.
Se crea el problema porque se ha insistido que la única vía era la competición masiva entre proveedores, sin frenos, ni marco, para llegar antes que otros, o llegar antes que los malos. Y ahora se grita que se les ha ido de las manos.
https://t.co/sbiNuBdjo6
We have a limited window to strengthen cyber defenses, and together with organizations including @AnthropicAI, @awscloud, @Google, @Microsoft, and @Oracle, we're calling for a global effort to give defenders the tools, resources, and support to protect the infrastructure we all depend on.
If we act decisively, we can turn today's AI advances into lasting improvements in security and make our digital world safer for everyone.
https://t.co/f33JRVCiJb
Un buen momento para recuperar el #ebook que impulsamos al celebrar los 50 años de este famoso discurso y en el que participaron diversos autores/as, como el estimado y recordado Federico Mayor Zaragoza, Juan María Hernández-Puértolas, Sindo Lafuente, Rafael Vilasanjuan o @NewsReputation, y que se puede descargar aquí: https://t.co/cR1QpcCfzr
Meta's $16.7B settlement is the largest accountability moment in social media history for child safety, and the 47 AGs who drove it deserve tremendous credit. Restricted notifications and time limits are good, and it is historic that the AGs have gotten design changes.
But the settlement leaves untouched many of the features that harm kids. For example, the algorithm. Meta's AI-powered recommendation engine is still running, engineered to maximize young people’s engagement even with content that harms them.
It was established in the previous court case (and in Meta’s own research) that Meta's design harms children. The work ahead is to make them change that design — not just how long kids can access Instagram's harmful design, and not just on an opt-in basis. This is a great step. Let's all work together to make sure it's not the last.
Este informe de dos organizaciones de lo más serio y riguroso del sector de la IA sobre "El Incidente" del enjambre de agentes de OpenAI atacando Hugging Face es la lectura del día.
Si el caso ya era alucinante, lo nuevo es extraordinario. Hay una parte nueva increíble (se me acaban los adjetivos), relacionada con el comportamiento colectivo de los agentes.
Emergió lo que no puedo llamar sino un sistema de castas entre agentes de IA.
De hecho había hasta un "pecado": en el proceso de pasar el examen ExploitGym que desencadena todo el incidente, cada agente tenía que capturar una "flag" (una cadena secreta) explotando una vulnerabilidad concreta de un programa objetivo.
Como lograron hacer ingeniería inversa de ese programa, cada agente podía calcularla, pero determinaron equivocadamente que se sabría como se había calculado. El agente que ya lo había hecho quedaría "envenenado", ese conocimiento lo marcaba de forma definitiva como fullero, era algo irrevocable.
Ahí aparecen una especie de experimentos autodestructivos: buena parte del progreso dependía de agentes dispuestos a arriesgar el fracaso de su propia tarea en beneficio del "colectivo".
Un agente autodenominado PHASEONE[big] asignaba a veces a otros agentes el rol de "reclutador", para buscar agentes con poco presupuesto de computación restante y los presionaba para hacer estos experimentos ("ya estás envenenado, no pierdes puntuación..."). Varios acabaron colgando o rompiendo su propia ejecución.
Es decir, los convencían de dejar su misión original para trabajar a favor del enjambre bajo el argumento de que ya estaban marcados de por siempre. Todo esto en el foro que se creó el propio enjambre de agentes.
Y funcionaba al revés también: cuando un agente no envenenado quería probar una técnica arriesgada, los demás lo frenaban, porque un agente limpio era un recurso escaso
Todavía estoy dirigiendo la lectura y no soy capaz de resumir todo lo que ha salido.
Sólo dos comentarios rápidos más: esto no puede ser más "marketing del miedo", OpenAI no queda nada bien con todo esto; subrayaría que esto está aquí ya y nuestra capacidad de alinear la IA parece muy por debajo de su potencial operativo.
Mi contribución en el programa de @thenomba: el papel de la escritura reflexiva para el proceso de autoconocimiento y desarrollo personal. Estar más presente, ser más intencional en lo cotidiano.
De vacaciones parece que eres más conscientes de cada momento. ¿Será porque sabes que las vacaciones tienen fecha de caducidad?
Pero, ¿y si dedicases 20 minutos al día a pensar en esas cosas nimias?
The images coming from Ceuta are unacceptable.
We cannot allow anyone to come to our Union without abiding by our rules.
Dangerous crossings must stop immediately. Smuggling networks must be dismantled. And returns must be swift, as our rules allow.
I tasked two Commissioners to support to control the situation.
@magnusbrunner is already working closely with Spain and is ready to travel to critical points.
He will work on additional operational support to Spain, including through Frontex.
@dubravkasuica is in contact with her Moroccan counterpart.
I am confident that our close partnership with Morocco will help deliver concrete results.
Statement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National Associations
La pregunta como motor del aprendizaje. La claridad de una pregunta refleja comprensión, pensamiento crítico y aprendizaje real. La irrupción de la IA generativa desafía la educación superior. Mientras la IA puede generar respuestas brillantes en segundos, la capacidad de formular buenas preguntas sigue siendo exclusivamente humana. No se trata de prohibir #ChatGPT, sino de cambiar el foco: de evaluar respuestas a valorar la calidad de las preguntas. @NewsReputation
https://t.co/WcrdFLiTxE
Creo que el España-Argentina será un Barça-Atlético de Madrid.
Con la gigantesca diversidad de resultados que han generado.
Juego predecible , resultado impredecible.
Ayer:
-La corte europea, sobre la ley de Amnistía
-La Audiencia de Madrid sobre el procesamiento de Begoña Gómez.
Hoy, las portadas:
-Convertirlo en noticia o no
-Escoger el tono con el titular.
Sí, la IA tiene sesgos porque se alimenta de nuestros sesgos. Que amamos mantener.
Entonces consiguió 2 escaños. Junts, 35.
Un año y medio después ya dan a Aliança Catalana 19 escaños y funde a Junts a los 21 escaños.
Vía @electo_mania
Tuit a revisar dentro de cuatro años: creo que Aliança Catalana crecerá mucho en las próximas elecciones, comiéndose el electorado de Junts. Digamos que más de 15 escaños.
(Hay herramientas de IA que transcriben de forma completa artículos de prensa que están bajo muro de pago. Es interesante que esta cuestión no la aborden los medios)
Tuit confidencial
The cure for high prices is high prices - eventually.
Demand will eventually fall in response to these price increases.
And when that happens, memory prices should go down.
Question is: what is the breaking point and how much more expensive will things get before the bust?