Opus 5.5 now leads our conceptual reasoning benchmark by a reasonable margin! (Better than the latest Fable and Astra)
Our team's very rough first impression from looking at its responses to a new eval under development is in line with these results.
Additional inference-time compute doesn’t seem to help frontier models much with our conceptual reasoning benchmark, LMCA. This is despite larger models continuing to perform better on LMCA. It’s also in contrast to what we see with LLMs solving coding and math problems where inference-time compute seems very helpful.
Some speculation for what might be going on: Use of inference-time compute is trained during post-training, and post training is largely focused on coding, math, and other domains with verifiable rewards. So plausibly models aren’t very effectively trained to use inference time on “fuzzy” tasks like conceptual reasoning. (By contrast, more effective use of pretraining data presumably does improve conceptual reasoning, because some of the pretraining data is conceptual.)
See the below inference-time scaling curves from us comparing the performance of models using different effort levels. You can tell that the models in fact think for longer with higher effort (cost goes up), it just doesn’t result in better performance. (For better readability, see separate plots for the different model families in thread, also includes a plot for Gemini models.)
(The black line shows the cost-performance pareto frontier: the best performance achievable at a given average cost, taking into account the possibility of randomising between different models.)
New CRI results are in! Highlights:
- Fable 5.1 leads but Astra 6 is really close.
- The gap between Anthropic and OpenAI models and the rest of the field is huge. The best non-Anthropic, non-OpenAI model ranks 11th overall.
- Meta is now the third-best lab on our metrics, pulling ahead of Google DeepMind.
- Gemini 3.8 flash still looks better on the CRI compared to Artificial Analysis.
(Full results on https://t.co/Fd3TkFYuhD)
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?"
Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this.
Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality.
This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below.
Official leaderboard website which we'll keep up-to-date: https://t.co/bsDF0uqHmw
Yesterday, we released the Conceptual Reasoning Index (CRI). The benchmark that we give most weight in the CRI is LMCA (Language Model Conceptual Argumentation). Our team of conceptual researchers poured 1000+ hours into generating its data.
We take position texts, arguments against these position texts, and then painstakingly hand rate the quality of the arguments according to a detailed rubric. We use those ratings to measure how good models are at rating arguments. We focus on argument evaluation because even on questions where we have little access to ground truth, such as on AI risk and philosophy, there is often much more agreement on which arguments are good.
We compare human ratings against each other for validation and to estimate ceiling performance. We find that human inter-rater agreement is much higher than agreement between models and humans. You can also read more about our validation process on our website (https://t.co/oOV5RGcopx). You can see in the graph that LMCA scores also correlate nicely with other capability scores.
You can also use the LMCA dataset to evaluate how well models can *write* arguments against position texts in our dataset, via an LLM grader prompted with human ratings of other arguments against the same position text. We will write more about this in the future.
Apply for dataset access here: https://t.co/ZgGpv866Ww
Mid-development arXiv paper (working on an update): https://t.co/C2cB53wCEz
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?"
Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this.
Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality.
This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below.
Official leaderboard website which we'll keep up-to-date: https://t.co/bsDF0uqHmw
@Mjreard As Chi notes, this isn't an example question -- sorry for the confusion! One can find actual example questions on the pages for the component benchmarks. E.g., here's the page for the highest weight benchmark: https://t.co/IyE6sDCxFm
I do think that there is transfer from general capabilities (and pretraining in particular) to conceptual reasoning. But I think that it's valuable to have better capabilities to help with safety work earlier (relative to other capabilities), and I think conceptual reasoning is currently pretty underelicited so that there are lots of low hanging fruit.