@oleg_murk very much agree! hosting a pragmatist manifesto meetup to get some conversations going in the research community if you have time to swing by: https://t.co/Pr9oTkIVCR
as a field our understanding of what rl is actually doing to models is still surprisingly rudimentary.
reward hacking is the obvious symptom: models exploit the gap between intended objective and measured reward, but we have weak tools for identifying the learned strategy itself, how it changes through training, and whether it survives distribution shift or evaluator awareness.
the core methodological problem is isolation vs realism. toy environments give causal control but often suppress the behaviors we care about; realistic environments preserve them but make attribution extremely hard.
i think the direction should be toward building an experimental science of objective formation under rl: checkpointing throughout training, perturbing reward structure / evaluator awareness / environment structure, then tracing when proxy strategies first emerge, what causes them to become stable, and whether the same internal mechanism survives as the environment becomes more realistic.
ideally this gives us something closer to a phase diagram of rl behavior: when reward hacking emerges, what training conditions produce it, and which learned strategies are genuinely invariant across environments rather than artifacts of a particular eval.
we should eventually be able to predict reward hacking before it becomes behaviorally obvious, distinguish genuine objective change from shallow policy adaptation, and understand which learned strategies remain stable across environments
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?
This isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training on user data* with very different privacy/IP implications. Sadly, AI cos don't like to disclose what they're doing.
- pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
- use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
- use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
"De-identification" is weak -- you can identify someone with a small number of bits, and long traces have more than enough. And it doesn't affect IP leakage concerns.
i wonder if it’s less about reward sharing and more about the combination of partial observability and trajectory-level optimization. once subagents each have only a local view, the optimal policy is to externalize intermediate state into a shared workspace. the message board is then a learned scratchpad for distributed inference rather than “altruism” per se
agreed that one important failure mode in verifier/task generation is mistaking difficulty for signal.
a lot of “hard” tasks are just poorly constructed. they introduce noise, ambiguity, or arbitrary complexity instead of isolating the capability we actually want to measure.
i think the missing piece is that most task generation lacks an explicit objective. before generating anything, we should be able to state:
• what capability are we testing?
• what specific failure mode are we targeting?
• what evidence would convince us the model improved?
then every task should be intentionally composed to maximize signal for that objective, not just maximize difficulty.
one direction we’re excited about is treating tasks as compositions of atomic capabilities, planning, abstraction, memory, retrieval, tool use, causal reasoning, etc. but the hard part isn’t composing them, it’s composing them while preserving what you’re actually measuring.
for example, if you’re evaluating long-horizon planning, adding difficult retrieval or ambiguous instructions can easily dominate the failure. a model might fail because it forgot a fact, misunderstood the prompt, or picked the wrong tool, not because its planning was weak. the planning signal gets washed out.
ideally, composition should preserve capability attribution. every added component should either (1) be controlled so it’s unlikely to fail, (2) have independent verifiers that tell you whether that component failed, or (3) be generated from a causal graph or dependency structure where you know exactly which capabilities each subtask depends on. otherwise, once a composed task fails, you no longer know why.
benchmark generation should optimize for information gain, not just hardness. the goal isn’t to make models fail, it’s to produce failures that are interpretable enough to drive the next training iteration
it's insane to me how many people take benchmark scores at face value. i spent 2 minutes looking at GDP-val tasks and realized there's a verifier item that checks for:
" The Word document includes the warehouse phone number 560-555-3867 (accepts common US formats such as 560-555-3867, (560) 555-3867, or 560-555.3867)."
but the prompt has:
"Include the address and phone number for Gravon Shoes’ warehouse, which is 555 Waters Avenue, Austin, TX 78726, phone number 455-864-3867."
the lack of good benchmark automated verifiers is jarring*, and most code benchmarks tend to lack the same verifier quality. this isn't to say GDP-val isn't useful or decent, but stuff like this is there across the benchmark if you dig deep. the "good enough" bar for data is finally starting to catch up to model performance. we need higher quality data, we need higher quality benchmarks to track improvements. my biggest peeve with how people construct tasks is that they mistake difficulty for signal. just b/c your benchmark has a 10% pass rate, doesn't mean it's good, and often, to make a difficult benchmark, people jump through loops. it's absurd that labs aim to hill-climb all these benchmarks but many have issues that would bar a model from getting better, so either you corrupt your model to improve or buy a bunch of data that doesn't actually help and look bad on your team. we saw this with OAI's paper on swe-bench, then subsequently swe-bench-pro, and im sure we'll see even more "benchmark debunks" soon.
making good benchmarks is hard, models have saturated most of the easy stuff and indie benchmark developers or contractors are no longer the best way of approaching it, because you can't have a hodgepodge task that a model will fail on. verifiers MUST be carefully constructed, the worlds the models are playing in must be realistic and without leakage, context has to be carefully selected. we're at a point now where 99% (arbitrary number) are saturated, and it's not "experts" but real expert engineers who understand what makes good training data.
*i understand that GDP-val specifically is graded by human experts, but for anyone running the tasks at home, it is still nuts to put out verifiably incorrect rubrics especially because others will use it.
every conversation I have had with a researcher from AmazonAGI has left me with hope for the future of computer use
If you are impacted or know someone that is impacted by this, I’d love to help, be it finding the right role at fleet or at any of the lovely labs & startups I know
guys it's not that deep. we just built a sandbox within 2 days. just open sourced it.
38ms startup (hot pool, p50 ~60ms p95)
1s (cold boot, p95, 1.5s p95)
auto resumes, 60fps VNC, PTY terminal sessions, instant fork, persistent volumes
costs $0.009/hr
Daytona has gone closed source. This is important for sandboxes, not for Daytona but for @e2b, who is now the only serious open source sandbox out there.
I want to take a second as a competitor and call out @mlejva for actually staying open source. It's not easy in the age of AI attackers.
If you're looking for the most powerful sandboxes, you know where to go (https://t.co/qT94Px2eGT). But if you are looking for the open source sandbox platform, there is now only one option: https://t.co/TtKIrCUz53.