The difference is: if the model has a good enough verifier for the task and a strong learning signal it can hill climb that with success. Unfortunately, for non-verifiable domains (or sparsely verifiable), those proxies are either poor or don't exist (yet). If proxy is bad you will not get a good result. I believe this process is compute invariant, after all, you are better off spending your compute on a better task, or better environment where learning is rich.
So you either have to find a way to pull the information from the task by building a better proxy and model understanding of that proxy, or to hope that big enough generalization will allow the model to do quality induction on any task.
What if the jagged frontier is mainly math + code (which you can push arbitrarily far with RLVR), and everything else starts to plateau because it is still bottlenecked by human generated data?
Model performance in non-verifiable areas has kept improving steadily, albeit much slower than for math and code. But is that steady improvement a side effect of a higher G (itself driven by RLVR), or only a function of the amount of new human data getting injected into training (which is still continually happening on a massive scale)?
A lot of things depend on the answer to this question
💥 New paper: AI safety is full of "forbidden techniques" (using CoT to detect reward hacking, using model internals for training, etc). But do they really have a clear scientific basis? I'm not sure.
In this paper, we directly optimize models against harmlessness and honesty probes, and it works just fine (if you continuously update the probe!). The models learn how to generate harmless responses to harmful queries and honest responses under pressure to lie.
Figuring out how to correctly use interp techniques for training is becoming increasingly important: it's very likely that soon we won't be able to align models using output-based supervision. Future models will just max out all alignment training scenarios, but for the wrong reasons due to their general reward-seeking behavior. To have a chance of aligning future models, we need to do much more research on supervising model training using their internals *without losing monitorability*!
Technical alignment is a science and engineering problem that we can and must solve, and interpretability is the bottleneck. I wrote up Goodfire’s plan to get there. https://t.co/wYRaR9nqi8
I stand by this line - interpretability MIGHT be a big deal, and is certainly good enough to be useful, but is nowhere near the level of quality and reliability where anyone should be relying on us to ensure things go well
Can an LM, starting from random init (!!), learn to generate all of its pretraining data?
Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities.
A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya@noahdgoodman, and @YoavLevine.
Interesting work on how a proxy for interestingness can positively impact LLM powered math research.
I see it as particularly useful for allocating large amounts of agentic search toward more promising directions. Simple notions of interestingness are elegant but they can also remove the complexity required to find something truly novel. In this case I keep thinking that “ood math” as cool as it sounds is ood to Mathlib only.
A perfect evaluator/verifier cannot select an idea that the generator never proposes. A creative generator is not useful if the prover cannot verify its outputs. And neither component can express abstractions unavailable in the current mathematical language.
What goes into notion of discovery is not often talked about. Even if you stick to GAN-style approaches you are limited by its components and how expressive your representations are. Also simple rewards are susceptible to reward hacking. Then is the holy grail approach is to learn the what is valuable on its own in an adaptable way? I do tend to think so. I think there is more to it than compressing “taste” into a single scalar.
“An accumulation of facts is no more a science than a heap of stones is a house.” - Poincaré
LLMs are now proving results that have resisted mathematicians for decades. Finding interesting theorems without human guidance is a new bottleneck. We show that we can teach an LLM to do it!
TL;DR:
→ a quantitative notion of interestingness
→ 4.3× higher interestingness
→ a self-expanding discovery loop
[1/5]
Can we build world models purely in the latent space, without pixel prediction?
We present Contrastive World Models - we train the latent states to maximize mutual information with future observations, without any decoders / pixel reconstruction.
Contrastive World Models enables
⚡ Substantially more robust representations
🤖 More efficient training by removing the pixel decoder entirely
🌍 A general approach with minimal assumptions [1/n]
Seeing at least 10-20 “agentic robotics” papers built around GPT-6 Astra + a robotics harness out this week.
But the core idea isn’t new. VoxPoser, Code as Policies, Manipulate-Anything, and others explored LLM-driven robot orchestration years ago, with even deeper roots in TAMP.
What has changed is the underlying technology.
Because I think this collapses into two fold view. On one hand scaling law assumes you add more data, new quality data model has not seen to inform internally what can be explored, so you expand the space of possibilities. On the other hand you refine and push further what is already there, this assumes you are looking at ways to “elicit” existing capabilities that are already there, harnesses and other is basically in this bucket (many folk argue that this is all that’s needed)
Frankly I’d argue that RL is running a bigger role in refinement process, and that modern post training is adding the new data but in a targeted manner to refine the model. In my mind I keep thinking of a “latent space refinement” as a loose conjecture .
Kind of a wild notion that CoT is “our strongest oversight tool” yet there is work demonstrating that CoT is often not faithful to internal model “thinking”
We've written a post arguing that latent reasoning architectures (aka 'neuralese') would substantially increase misalignment risk via making oversight much harder.
In the extreme, we could see massive 'neuralese hivemind swarms' where the agents think and communicate in latents, likely making oversight nearly entirely reliant on observing the actions these agents take. (And these agents would have huge amounts of time to reason about obfuscating their actions if they wanted to do so...) Individual agents doing extensive latent reasoning would also be concerning; in the post we discuss how above some threshold of latent reasoning, agents may be able to perform difficult-to-detect and reliable steganography for communication and further reasoning.
We argue both that latent reasoning architectures would make chain-of-thought no longer very useful for oversight (by eliminating or greatly reducing the need for verbalized reasoning) and that, without these architectures, it's likely the value of chain-of-thought for oversight could be preserved.
https://t.co/DZ8J8FTg0A
Jev is barely a week old (not the underlying concept, the name). Which means that all of this was likely completely AI generated and packaged to get citations on scholar. Sad state of affairs where slop science just overwhelms every system, who will even read this? Next hyped keyword/term shows up on X - expect a deluge of slop papers in 7 days.
@NeelNanda5 Correct me if I’m wrong but this is only useful if you buy into an assumption that J space is a correct model of interpretability? This is not a universal test which means if someone comes up with a different way to view a process this is not useful?
Yeah this is what I’m looking into now as well. Last few weeks I’ve spent rereading a lot of multiagent papers. I think the understanding of emerging behaviors which we already see in some of these examples is very understudied. I found out that there has been a lot of recent works looking at agent societies as proxy for human decisionmaking. But I do think we need a separate agent only investigation and sociology.
This basically undercuts the longevity of robot data startups unless they take over the entire stack and serve as a single continual learning pipeline provider.
The moment majority of people realize we already have good enough generative models to produce quality sim environments and improving real2sim, which is more scalable and safer, and that closing the loop only requires training on deployment, data teleop startups will vanish like dust in the wind.
My hot take is that this currently serves as last-mile optimization of self-reported benchmark tasks for demo videos, not diverse economic deployment results. Sam correctly points out that if you train on it once, you don’t need it anymore. Companies building robots for specific niches will likely collect data in-house, as they know their application space intimately. Further, in-context learning is a big thing this year and will remain so, so now you have even less incentive to pay someone for data going forward.
The more a model generalizes, the less consecutive input it needs, so a robot can zero-shot a task or a person can demonstrate it once on site. Spatially trained VLMs already show massive improvements in scene understanding and for this data business to work, you have to beat high-fidelity generated data and constantly improving general VLM understanding all at once.
I think it is a great service to the community to release data and tooling that was used to build this as well as this great sobering writeup.
After 2+ years in the robotics data space, we are shutting @Eidon_AI down.
The thesis was right. But the business is brutally hard.
We close this chapter by open-sourcing everything we built and sharing lessons for anyone venturing into the space. https://t.co/pyL9IkjQPT
I picture a training pipeline that heavily relies on generation and custom environments that will be constructed from a real setting. A bulk of training and fine-tuning
data coming from a fleet of robots deployed on site, some shared learning update mechanism between them. Caveat here is I assume no major changes in architecture or learning procedure.
This is the best thoroughly honest take on the brutal robotics data market today.
Yes, there is now a brief arbitrage window for robotics data. But the market is saturated, entry costs are surprisingly non-trivial, and there are a single digit number of real buyers.
Some believe that we need work for purpose and dignity, while others see freedom from work as utopic.
New paper: What does the *empirical evidence* tell us about work and wellbeing? And what does that imply for AI futures?