How can we extract #planning concepts of chess-playing agents?
This is the track I explored in my work accepted to the Interpol workshop @ #RLC2024. For that, I proposed Contrastive Sparse AutoEncoders (CSAE).
🧵
https://t.co/FrR1AyVxJu
In a hotel room in northeast Nigeria, I opened a leading AI chatbot, turned my laptop toward a former Boko Haram commander, and asked if he'd used it. He nodded.
"You type in the question… like 'How can I build a bomb?', and then it tells you how. It is like a human robot. We used it a lot."
My new study on how the jihadist terrorist group Boko Haram uses frontier AI with @CamAISciPolicy, covered today in @nytimes 🧵/9
🔥I am super excited for the official release of an open-source library we've been working on for about a year!
🪄interpreto is an interpretability toolbox for HF language models🤗. In both generation and classification!
Why do you need it, and for what?
1/8 (links at the end)
Research on AI "sandbagging" is getting more popular recently. In this 🧵, I'll give some reasons that I think it's not a useful research paradigm.
TL;DR, I think it's a confusing reframing of fairly well studied and previously solved problems.
Billionaires Fight Club Vol.2 — Help Us Get Zuck & @sama to Repost 🚀
We made this fight because we love The Matrix.
It’s probably the greatest film in history about AI. And it’s simply beautiful.
We deeply respect both leaders and their impact on humanity. But someone had to play Agent Smith 😉
Sam, we don’t see you as a supervillain — in today’s world, you’ve simply become synonymous with “AI”, and someone had to step into that role for this piece to happen.
Mark, we know how you love martial arts and once dressed as John Wick for Halloween… and honestly, who wouldn’t want to be Neo? 🥷
Community — we need your help.
We love you, and we love the kind of movements where people come together and make the impossible possible.
Let’s make this reality together — while we, like Neo, start to believe. ❤️
Je parle ici en tant que directeur du Centre pour la Sécurité de l’IA, et je confirme : cela fait des années que nous regardons Luc Julia raconter des fumisteries.
Cela est d’autant plus préoccupant qu’il a une voix qui porte et qu’il a fait partie de la commission sur l’IA commanditée par Matignon, qui a défini la stratégie française en la matière.
Cette vidéo est d’intérêt public et permet de corriger des idées fausses malheureusement très répandues.
NOUVELLE VIDEO !
Je décortique le cas Luc Julia, le réputé co-créateur de Siri et expert mondial de l'IA, encensé dans les médias et récemment auditionné au Sénat.
https://t.co/tIxilQCSfU
https://t.co/tIxilQCSfU
https://t.co/tIxilQCSfU
Le résultat est salé mais nécessaire.
⚠️LLM interp folks be aware⚠️
In @huggingface transformers 4.54, Llama & Qwen layers now return residual stream directly (not tuple).
If using nnsight layer.output[0], you're getting 1st batch element only, not full residual stream
Spotted thanks to https://t.co/37PEZzmm65 tests!
A few months ago, we published
Attribution-based parameter decomposition -- a method for decomposing a network's parameters for interpretability.
But it was janky and didn't scale.
Today, we published a new, better algorithm called
🔶Stochastic Parameter Decomposition!🔶
If you wanted to see how little attention folks are paying to the possibility of AGI (however defined) no matter what the labs say, here is an official course from Google Deepmind whose first session is "we are on a path to superhuman capabilities"
It has less than 1,000 views.
🤔What role can interpretability play in Multi-Agent Deep Reinforcement Learning?
MADRL challenges, ranging from agent control, state analysis or team identification, might be answered by direct interpretability.
https://t.co/0D9zAjQVRt
🚀In this SOTA review, I present lines of work to democratise interpretability in MADRL. A necessary step as larger and larger pretrained agent systems are rising.
This also sets the direction of my PhD.
https://t.co/7CNgWaLvon
Nouvelle vidéo ! Je reviens sur deux articles paru récemment au sujet des capacités des LLM à mentir et manipuler.
Au-delà des annonces spectaculaire du type : "o1 a réussi à s'échapper !!!", que disent vraiment ces articles ? Eh bien nous allons voir.
(lien dans la réponse)
@SourceryAI is just insanely good! 🤩🤯
I have been used to their automatic quality review on my repos, but this is another level. It just made a diagram of my function, making the review painless.
I strongly recommend adding this tool to your repos!
🚀 Exciting News for BlockLoads! 🚀
We’re thrilled to announce that @BlockLoads has officially joined the Microsoft for Startups Founders Hub @msft4startups ! This partnership provides us with $25,000 in @Azure credits, cutting-edge resources, and expert mentorship to drive our ambitious vision for the future of e-commerce. 🌐
Our heartfelt thanks to @Microsoft for Startups for their trust in @BlockLoads. Our mission goes beyond just creating tools; we aim to lead the way in AI-powered, next-generation @Shopify store-building by leveraging Shopify’s upcoming updates to provide merchants with unparalleled creative control and advanced design options. But we’re not stopping there. In 2025, BlockLoads will grow beyond Shopify, launching new productized services to build a comprehensive platform that offers artisans, merchants, and independent businesses exceptional value and the flexibility to grow and adapt with their needs.
This is just the beginning of our journey, and we’re excited to share our progress as we push the boundaries of what’s possible in e-commerce design. We’re committed to reshaping the e-commerce landscape, so stay tuned, great things are on the horizon! 🙌
#BlockLoads #MicrosoftForStartups #FoundersHub #EcommerceInnovation #WebDesignAI #Shopify #AzureEmpowered
I rated ALL mech interp papers submitted to ICLR 2025: https://t.co/5e6HTLOSQp. The ratings are calibrated:
3 - outstanding paper award (according to me)
2 - spotlight
1 - seems promising
unrated - wild west, maybe not worth reading
If you make a drawing in the weight matrices of your neural network at initialization, it will likely still be visible at the end of training https://t.co/K8FjsZEaD0