We just updated CI-aware bench, which measures whether frontier models are aware of AI control interventions made to their text.
When we first talked about this in spring, most models were near chance, and some readers complained that the bench was unrealistic...
Astra nearly saturates it now.
For more details, check out Joachim's thread linked below:
I'll be in San Francisco all week for @COLM_conf and to talk about investing with @RockboundCap. Let me know if you want to chat about the strategies used in high net worth investing and how they can work for you.
"Sphere Encoder 2"
This paper fixes why one-step autoencoder generation looks blurry.
It trains the latent rotation all the way from an encoded image to the 90° equator region where random generation actually starts.
Then instead of forcing those highly ambiguous latents to reconstruct the exact source image pixel-by-pixel, it uses semantic, latent-consistency, and feature-space score-matching losses so the decoder can generate a sharp plausible image rather than average many possibilities into blur.
On ImageNet, it reaches FDr6 4.92 with just 4 steps and ~23x less sampling compute than PixNerd, though its gFID is worse.
https://t.co/blg7FN8Onz
Excited to share Pinocchio, our COLM 2026 paper! 📄🧵
Pinocchio is an external calibrator that estimates whether a black-box API model's answer is correct, in a single forward pass, with no logits, weights, or sampling. It transfers zero-shot to 13 models it never saw. Including Frontier models like GPT-6 Astra
Frontier LLMs like Claude don’t come with uncertainty estimates, and their verbalized confidence is bad. We trained Pinocchio, a lightweight model that assigns confidence to outputs of popular API models, making it fast and easy to get well-calibrated uncertainty estimates. 📄🧵
Frontier LLMs like Claude don’t come with uncertainty estimates, and their verbalized confidence is bad. We trained Pinocchio, a lightweight model that assigns confidence to outputs of popular API models, making it fast and easy to get well-calibrated uncertainty estimates. 📄🧵
1/9 🧵 What if agents should learn to predict what they will never generate?
ActObs supervises actions AND environment observations during SFT, teaching action consequences with no extra data, parameters, sequence tokens, or forward passes.
📄 https://t.co/jgdUO3Y6Yp
Making models deeper through recurrence in depth does make them harder to monitor
but his whole discussion seems very quaint/academic when OAI has deployed models for many months that reason like this
Generative models can’t discover what they can’t reach.
We’re excited to introduce ActFlow: a continued pre-training scheme that actively expands the valid design space reachable by flow and diffusion models. We call this generable set expansion — a new learning principle for out-of-distribution generative modeling, and a step toward evolvable search spaces for scientific discovery. (1/5)
New Anthropic Blog: We asked whether training a dedicated deception monitor gives you something beyond just prompting a capable model to detect deception. In-distribution it does, but OOD the advantage mostly disappears: fine-tuned detectors barely beat prompted baselines, and larger prompted models often outperform them. Read the full post here: https://t.co/O0ta1KnPL5
Thank you @FabienDRoger@rowankwang for mentoring this project !