AI (esp.., alignment/post-training for safety and improved reasoning), formal semantics, pragmatics, philosophy of language, mind and cognitive science
New preprint available! "Emergent alignment and the projectability of ethical personas" (joint work with Youngchan Lee, Calum McNamara, and Alejandro Pérez Carballo).
https://t.co/EZcEl1piMf
Here's a summary.
Recent work on "emergent misalignment" has shown that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. In this paper, we explore the converse phenomena—“does finetuning LLMs on narrow tasks result in broadly aligned behavior?" ---and use it to study important out-of-distribution characteristics of several well-known ethical systems (e.g., Consequentialist, Deontology, Virtue Ethics), when used as central components of AI alignment strategies.
Here are some of our main findings:
3. The specific ethical perspective used to fine-tune a model—even if applied to just a narrow safety subcategory—leaves a coherent and distinctive "moral fingerprint" on the model’s overall behavior. The resulting (narrow) models didn’t just become generically "safer"; instead each of them took on the specific “ethical persona” encoded in the constitution which we used to create the safety data used for finetuning.
Nice convergence with our work on pre-pretraining for language. Transformers need a lot of training tokens to acquire basic processing mechanisms. Pre-pretraining on procedurally generated synthetic data speeds up the process and improves data efficiency.
The 3rd edition of my book Deep Learning with Python is being printed right now, and will be in bookstores within 2 weeks. You can order it now from Amazon or from Manning.
This time, we're also releasing the whole thing as a 100% free website.
I don't care if it reduces book sales, I think it's the best deep learning intro around, and more people should be able to read it.
New blog post: Philosophy of mind has changed a lot, with a truly massive shift toward empirically-informed work — but we still have in place various norms and institutions that were created decades ago and don’t make sense for the field as it is now
BOOM! New review paper by Eleonore Neufeld et al. reviews a host of recent experimental studies on generics…
…and argues that there ‘s growing evidence against the whole idea of a special connection between generics and social/political beliefs
https://t.co/sQTtP8rSsv
Just published open-access in the Journal of Semantics: Bernard and Champollion, "Negative events and compositional semantics". https://t.co/kKlJ7TIgZ8