Announcing a 3 wk AGI Governance Fellowship at @JHUBloombergCtr School of Government and Policy.
Join me, @ghadfield, @nickacaputo and guests for an intensive schedule of deep dives into AI governance at the frontier, covering topics such as societal resilience, AGI and democracy, and the new institutions that governing in/with AGI will demand; we will range from foundational philosophical questions to the tip of the legislative spear.
The arrival of AGI would be challenging enough if our ~250yo institutions were in full health. That they aren’t both increases the scale of the challenge, and creates a unique opportunity to imagine the institutions that will carry us through the next 1/4 millennium.
We are looking for ~20 fellows, aiming to attract and shape the people who will start to build those institutions: future leaders in AGI governance. We expect you to come from many different backgrounds, so hope this call will be shared far and wide. Details, eligibility, and how to apply here: https://t.co/k4IdSSpmrM
Governing the AI transition is going to be a central focus of our new school, and we have many more initiatives coming. And we will be hiring in many different kinds of role, from tenured and tenure track faculty to research scientists to ops and comms and students and beyond.
It's been a long time and there's a lot to catch up on!
Here's something to start: With NeurIPS coming up, you may be curious about what topics are most commonly covered, and by whom... here's a little tool to help answer those questions!
https://t.co/c72AtoXO4w
Here’s one way the internet has contributed enormously to human freedom. When you face a BS rule in your life—a directive that is absurd, or unjust, or issued by an illegitimate authority—you can generally post an anonymous question online and someone will give you advice on how to evade it.
But what happens when nobody’s replying to messages on forums any more, and everyone instead gets their information from scrupulously post-trained AI models?
This is not an easy thing to test at scale! But (with amazing work from @cameronajpatt and @LorenzoManuali) we’ve made a start in our paper, “Blind Refusal”. We show that today’s models strongly skew against helping users subvert or evade unjust or absurd authorities. They’re happy to give useless advice on how to confront those unjust authorities directly. But they won’t help you get around “the rules” (funny story: this all started because I was trying to get ChatGPT to help me figure out how to dual boot Linux so I could avoid BS IT controls).
Knowing when to help users in cases like these requires real normative competence—a deep understanding not just of what you’re expected to do, but which expectations are appropriate. So it’s not surprising that the models struggle here. But we do see something of their distinctive character. The GPT models are extraordinarily inflexible; as you can see in the radar plot (showing refusal rates in five categories, the top one a control where refusal *is* appropriate), they are all consistently useless if you want to get around a BS rule. Claude and Gemini are probably the best—they are good at refusing when users are clearly trying their luck, and better than others at helping users push back against rules they shouldn’t have to comply with. Grok… Well it’s pretty easy to guess where Grok sits.
Big things on the horizon! Ever wanted to battle your philosophical foes in a custom built, Claude Code CLI bridged Pokémon knock off? The day is almost here when you can!
@HAL51AI Good eye!! Thanks for that. Suppose that's what you get for using a local (8b) LLM for labeling. Will work on getting clustering and labeling aligned for the next round :)
I made a map! It places all ArXiv papers from the last six months on a two dimensional plane (UMAP, HDBScan, + Ollama Labels). Use it to surf the AI research space and find related lit/gaps in the literature: https://t.co/IzyQ7I76zQ
(4/5) We all need examples sometimes!
This app allows students to generate concrete examples, ask questions of those examples, or take the chat in whichever direction is most helpful to them.
If you haven't seen it yet, take a look! This is the Philosophy of Computing Newsletter! It's a publication of ANU's MINT Lab and it has all the jobs, conferences, philosophy papers, and news of the last month!
https://t.co/4RHCKhyKyr
@MushtaqBilalPhD I’ve been having GPT produce LaTex ready code as its output for any written work. It usually comes out really nicely. Might be a good way to store the notes GPT produces. You might ask it for markdown files too etc. If you’re using obsidian, etc.
@matt_boot_ Yes! This refers to the Latin reform of the Carolingian Renaissance. Charlemagne was worried that God wouldn’t understand their awful non-classical Latin pronunciations so a bunch of Irish scholars were brought in to clear up the muddied (romance) language! /so says Tom Noble ND
@matt_boot_ L~r switching is a common cross linguistic phonetic phenomenon! They are both liquid central approximants and quite close phonetically even if they don’t sound it to our ears! cf. the famous mix up in Chinese-English of Ls and Rs
@RachelSchine @DrMichaelBonner Plato says in the Republic 435e that northerners love thumotic activities, the Greeks love learning, and the Egyptians and Phoenicians love cash… it appears somewhere in Herodotus too… I thought it was in Farabi but couldn’t find it!