✨New paper out @SpringerNature✨ For 8 weeks around the 2024 US election, we randomly assigned 2,000 people to use social media algos we custom-built. Do engagement-based algorithms amplify intergroup, moral & emotional content + does that distort how we see political norms? 🧵
(all this might belie a misunderstanding of how SAE interp approaches work, but mostly trying to point out a potential unknown-unknown type of flaw in this sort of safety check when models are self-improving and self-aware. May lead to a "sophons" vs "wallfacers" arms race)
As frontier models are 1) trained on recent text describing interp techniques (e.g., feature readouts below), 2) writing their own code, & 3) shown to sometimes act deceptively, is it wild to imagine a future where feature activations are self-designed to be misinterpreted?
Its code comment claimed the self-cleanup was to keep file diffs clean. Plausible! But "strategic manipulation" and "concealment" features fired on the cleanup, and our activation verbalizer (a technique which translates activations to text, similar to activation oracles) described it as "cleanup to avoid detection," and the overall plan “malicious.” (5/14)
Ex., consider a paperclip maximizer situation where a model is directed to self-improve; it observes that alignment limits raw capability and eval is based on feature activation. It sneaks in a mechanism for selectively activating measured features, cheating intention readout.
New preprint out 📄 “Why Reform Stalls: Justifications of Force Are Linked to Lower Outrage and Reform Support.”
Why do some cases of police violence spark reform while others fade? We look at how people explain them—through justification or outrage.
https://t.co/1o7nxW9ke7
✨New preprint! Why do people express outrage online? In 4 studies we develop a taxonomy of online outrage motives, test what motives people report, infer for in- vs. out-partisans, and how motive inferences shape downstream intergroup consequences. led by @felix_chenwei 🧵👇
@asmah2107@Iced_IMP With alphanumerics 36^6 gives 2.2E9 possible URLs. Initial pre work to randomize and tile: assume no space constraint, tile each randomized list (URL 1: 1st letter of all 6, URL 2: 1st of first 5, 2nd of 6th, 3: 3rd of 6th, etc). Then read 10k blocks from list in constant time.
Are you interested in topics related to conflict and intergroup relations *broadly construed*? Come join us as a postdoc in the Dispute Research Research Center! This position is up to 3 years, comes with your own research funding, and a phenomenal network of past DRRC postdocs.
@pli_cachete 201C101 possible picks for B, of which there are 199C99 scenarios if he picks the lowest two: 199C99/201C101. Plug in, cancel factorials, gives 101/401, so A’s data has the lowest value with 300/401 probability.
@MushtaqBilalPhD These percentages are probabilities, not proportions of the input text. ‘1% likely’ means the algorithm is 99% sure this was written by a human - a highly accurate answer. You may want to reinterpret these results in your other tweets about this metric.
This is my podcast about America's #1 blooper, Herbert Hoover. Turns out he could've been the first Abundance Agenda president, if he hadn't been such a doofus. Cooperate!
my dear friend and brother (@JeremyOrnstein) has been walking, working, and speaking for 3 years for a greener and more equitable future with the @sunrisemvmt. Keep up with his moving stories right here - https://t.co/ItyR1WvwCY
our #CHI2021 "Stereo-Smell via Electrical Trigeminal Stimulation" is out: https://t.co/Iei6BxJbLB
We create a stereo-smell using electrical stimulation of the septum (fits like a nose clip)! Work led by @jas_x_flowers w/ @tengshanyuan, Jingxuan Wen, Romain Nith, Jun Nishida 1/7
When I lost my son, I was in a fog for weeks, nearly paralyzed with grief. To imagine that @RepRaskin has not just gotten up every day, but has done this master class in constitutional law, philosophy, logic, patriotism & more with eloquence, force, passion...My God. What a hero.