New paper from @Preethi__S_ internship @cohere! https://t.co/7jNYXBuOQv
LLMs are very commonly used in HR, but most fairness work on it in a generative context is contrived and unrealistic. We build fairness tests based on observed real usage with super interesting results 🧵
is anyone else also constantly scrolling on their endless Claude/ChatGPT thread? yo @OpenAI@AnthropicAI a file outline we can click on to get back to previous questions would be great
Do LLMs plan the content of a paragraph at its onset? Together with Nicholas Pochinkov, Angelo Benoit,
@lovkushatleeds, and Zainab Ali Majid, we take a mechanistic interpretability approach to investigate this research question.
https://t.co/V8GL4lFxoL
Our findings:
By transferring the activations from the final layer (amongst others), we are hinting at what the next token is. We thus added 'cheat' tokens to the neutral baseline.
(2) neutral + 2 tokens compares with our transferred generations wrt closeness to original
My first paper in AI Safety now on arxiv. https://t.co/OU4Qahw15m
Big thanks to my teammates and to the SPAR orga using team. Huge credit goes to Lucile for doing the heavy lifting with writing and logistics for the paper.
started @BlueDotImpact's AI Governance course today! Really impressed with the quality and organisation. Excited to learn about policies that improve AI safety
(1/2) as a side hustle: I've decided to start a series, writing reviews on Mechanistic Interpretability papers. I'm mainly going to sample from @NeelNanda5's v2 list
mech interp is a subject i've been getting super enthusiastic about in 2024 since discovering @ch402's work
takeaways: I'm learning a lot!
This project touches upon LLM-specific challenges, which I have been reading about but have not yet had hands-on experience with. I'm very excited to be part of this
takeaways: I'm learning a lot!
This project touches upon LLM-specific challenges, which I have been reading about but have not yet had hands-on experience with. I'm very excited to be part of this
PhD tip: there are plenty of mentorship programs for you to learn about topics outside of your lab's expertise!
I got accepted to SPAR & am now working on an AI Alignment project looking at how LLMs plan for future tokens, supervised by Nicky Pochinkov https://t.co/GanC3kY7K8
about the project: we're investigating how LLM "plan ahead" i.e. how current-token hidden states encode data directly useful for possible future tokens. Our aim is to do so in a compute-efficient way
On the heels of Open AIs letter in safety - we need a global governance body for generative AI systems. Here’s my thoughts on how we can build one.
https://t.co/FDMpv1zzFl