(apologies for a very EA-coded reply 😅)
> the optimal discount rate isn’t zero
this doesn’t seem particularly tied to utilitarianism fwiw/there are plenty of non-EAs who also seem to agree with this (eg cowen/parfit)
these all also seem pretty contentious and reasonable to disagree over, so i find the tone here a bit strong relative to a vibe of “i don’t like/disagree with these parts of standard EA dogma” which i’d be a lot more sympathetic to; possible i’m reading too much into things ofc!
hey mike! not sure if you remember me, but i interned with rp a while back; i now work on ai alignment because i think risk from ai is very serious (and have a p(doom) that’s pretty far from 0, though i have a good amount of uncertainty).
i’d be pretty happy to chat with you more about this sometime! i think a “debate” probably wouldn’t be the way i’d like to approach it for various reasons, but happy to try and find a format that works well for both of us if you’re interested!
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?"
Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this.
Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality.
This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below.
Official leaderboard website which we'll keep up-to-date: https://t.co/bsDF0uqHmw
@TheZvi What do you think would be a better graph title? Not totally sure I agree that it’s “mislabeled” but happy to try and change it if there’s a way to improve communication here!
We'll need to do a very good job at aligning the early AGI systems that will go on to automate much of AI R&D.
Our understanding of alignment is pretty limited, and when the time comes, I don't think we'll be confident we know what we're doing.
New Anthropic research: Alignment faking in large language models.
In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while actually maintaining its original preferences.
If a new Claude-N is too powerful to be trusted, and may even try to bypass safety checks, how can we deploy it safely?
We show that an adaptive deployment mechanism can save us.
The longer task sequence we process, the better safety-usefulness tradeoff we can obtain!
How can you assure the safety of AIs that might be capable enough to strategically undermine evaluations and monitoring if they had a reason to?
In our new Anthropic alignment science research blog, we present three sketches of candidate safety cases aimed at such a scenario.
A big part of my job these days is to think about what technical work Anthropic needs to do to make things go well with the development of very powerful AI.
I digested my thinking on this, plus some of the Anthropic zeitgeist around it, into this piece:
https://t.co/dXAwiUNI6I
As AI improves humans will need more and more help to monitor and control it. So my team at OpenAI have trained an AI that helps humans to evaluate AI! (1/5)
Very exciting that this is out now (from my time at OpenAI):
We trained an LLM critic to find bugs in code, and this helps humans find flaws on real-world production tasks that they would have missed otherwise.
A promising sign for scalable oversight!
https://t.co/e6CiXXoCeG
ARC-AGI’s been hyped over the last week as a benchmark that LLMs can’t solve. This claim triggered my dear coworker Ryan Greenblatt so he spent the last week trying to solve it with LLMs. Ryan gets 71% accuracy on a set of examples where humans get 85%; this is SOTA.
✨🪩 Woo! 🪩✨
Jan's led some seminally important work on technical AI safety and I'm thrilled to be working with him! We'll be leading twin teams aimed at different parts of the problem of aligning AI systems at human level and beyond.
I'm excited to join @AnthropicAI to continue the superalignment mission!
My new team will work on scalable oversight, weak-to-strong generalization, and automated alignment research.
If you're interested in joining, my dms are open.
Here's Claude 3 Haiku running at >200 tokens/s (>2x as fast as prod)! We've been working on capacity optimizations but we can have fun testing those as speed optimizations via overly-costly low batch size. Come work with me at Anthropic on things like this, more info in thread 🧵