My very first co-first-authored paper was accepted to EMNLP 2026! ๐ฅณ
TL;DR: We observed a serial position effect in Mamba and took a peek at how it manifests internally in the model.
See you in Budapest! (Hopefully ๐ )๐ญ๐บ
Read our paper on scaling interpretable LLMs, we show that interpretable LLMs are not only possible but can scale both on typical generation benchmarks and interpretability benchmarks. https://t.co/4BNrqgbTlB๐
polisi akan ajari anak sekolah tentang AI:
- polisi lagi belajar soal AI
- biar nanti bisa ngajarin ke anak-anak sekolah
- ada 65 polisi yang dilatih dulu,
- biar mereka jadi "guru" buat polisi-polisi lain
- soalnya banyak anak sekolah udah pinter pake AI sendiri
- tapi enggak ada yang ajarin cara aman pakainya
- polisi mau ajarin 10 juta anak sekolah
Researchers proved every major LLM is secretly obsessed with Japan.
And they finally figured out why.
For years, weโve been told that AI is entirely Western-centric, that it just reflects Silicon Valley and American values.
A landmark paper by Cardiff and Basque researchers tested 31,680 cultural prompts across 24 languages on frontier models like ChatGPT, Claude, and Gemini.
The results shattered that assumption.
In six out of eight frontier models, Japan was the single most frequently referenced country when asked open-ended cultural questions.
Ask about traditional dances, festivals, or everyday practices in an open context, and the AI defaults to Japan.
Over and over again.
Here is the twist nobody expected.
This bias doesn't come from raw pre-training internet data.
The researchers tracked where the obsession forms. It emerges after pre-training, during the supervised fine-tuning and alignment phase when humans teach the AI how to behave.
Why Japan?
Because decades of global soft power, rich cultural export, and clean, universally admired digital archives make Japanese culture uniquely "safe" for AI safety filters to lean on.
When labs train models to be harmless and universally pleasing, the AI defaults to the cultural equivalent of comfort food.
It avoids controversy by talking about anime, sushi, and tradition.
There's a drama over a fan project getting attacked coz the dev uses an AI coding assistant
Crazy moral whiplash and misconception to the extreme of blanket "AI = bad" (doubt they know AI isn't just LLMs)
Huge disconnect on AI understanding among the general public and experts
Dulu kami kirim putra2 terbaik melalui Indonesia Mengajar dan Guru Garis Depan utk mengabdi ke pelosok sbg pengormatan thdp setiap jengkal wilayah NKRI.
Memberi stigma buruk pada Papua, Malut, dan daerah terluar NKRI sbg tujuan mutasi hukuman sama dgn menghukum daerahnya juga.
Did you know?
Pangram learns the difference between Claude, ChatGPT, and Gemini in its internal representations, even without being trained on it!
This signal is increasingly recoverable throughout the network, reaching 91% accuracy on a simple linear probe!
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. ๐งต