@yuntiandeng I feel like there’s a lot of anthropomorphism in the AI Safety and Alignment space. Especially in the blogs A\ has published regarding “Scheming”. I’m interested in exploring some of these questions in a academic setting to get open source and replicable results.
@MariusHobbhahn I think this makes it even more important for AI safety, because imagine how easy it would be to spread a particular political viewpoint and shape belief systems.
I wonder if we should test out the sweet spot of RL tuning so that it’s useful but still steerable.
Our position paper has been accepted to the Scientific Understanding of Foundation Models (Sci-FM) workshop at @COLM_conf.
It's time for our community to shift from post-hoc narratives toward genuinely exploring the science of models.
See you in SF this October! #COLM#NLProc
This paper resonates so much: https://t.co/pSvnW6b5Lz
We should really consider improving smaller models to mimic human writing. I have always felt that huge coding optimized model like fable can never truly write something magical like George R. R. Martin’s novels.
My students sometimes ask why they should memorise things in the age of Google and LLMs. But internalised facts are your bullshit filters and your raw material for creative association. Facts outside your head are inert.
Check out this very creative work from Kunal, exploring the effects of post-training alignment through the lens of physical crystallization! There is a lot that AI research should take from the natural sciences, including more metaphors
That's the predictive framework we're aiming for.
Paper: https://t.co/9shsVSo3YN
"Towards Physical Intuitions for Alignment Dynamics: A Case Study With Randomness Crystallization"
w/ the amazing @PeterWestTM & @universeinanegg
We're good at measuring what alignment does to LLMs. We're much worse at describing how it happens over time. Our new paper argues alignment research should borrow predictive theories from physical sciences; using crystallization as a case study. 🧵
Right now we mostly diagnose alignment after training. We want theories that predict it beforehand: which latent behaviors become the "seed"? Which vanish? Which are hard for RLHF/DPO to recover once SFT has crystallized the model?
Again, please do NOT confuse a model of knowledge with a system with intelligence, any longer. Intelligence is the ability to produce knowledge. Knowledge is never truly generalizable, no matter how much; Intelligence is always generalizable, though capacity maybe limited.
Presenting my #ACL2026 paper: "Simple Agents, Biased Judges: Efficient Multi-Party Dialogue Generation & The Evaluation Gap"
📄 https://t.co/NJE4hejbOX
I'll be in San Diego from 1-8 July.
Looking forward to chat about alignment & creativity in LLMs more broadly. #NLProc
Presenting my #ACL2026 paper: "Simple Agents, Biased Judges: Efficient Multi-Party Dialogue Generation & The Evaluation Gap"
📄 https://t.co/NJE4hejbOX
I'll be in San Diego from 1-8 July.
Looking forward to chat about alignment & creativity in LLMs more broadly. #NLProc