Training on reward-hackable tasks is the single biggest source of misalignment. Verification has to come from an independent third party, not the vendor, not the lab.
We are making acceptance testing as a service, independent verification layer, and a marketplace for RL tasks.
If you're a strong researcher who wants to build that verification layer and own it, DM me. We have a real wedge in the market and a clear view of what to build.
New research: Training a Misaligned Reward Seeker
What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable.
In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.
Read more: https://t.co/gs2ZjYkPan
A must-read for data vendors looking to enter the Korean market.
I also help vendors enter the Korean market. If you have data relevant to AAII, DM me.
Korea’s Trillion-Dollar Sovereign AI Investment: Nvidia Wins, Hynix Loses
Korea hosts a Squid Games, National AI Tournament,
the best non-Chinese open source model gets eliminated,
why Nvidia needs open source, implications for Hynix and Samsung
https://t.co/E3eiJx4VzN
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
https://t.co/Nb2un9oNJR
Top 10% at YC may not be a bad result for a pre-revenue, 3-month-old, no-name solo founder working out of a small room in Korea.
But for the next batch, I hope I’m too busy building to even have time to apply.
Over the past three months, I received inbound interest from 5+ top Silicon Valley investors, a Fortune 500 AI VP, and Harvard/Stanford professors. It’s been a reminder that if you have a clear vision and strong research, the barriers to the world may not be as high as they seem.
I’m looking for a cofounder. Someone who wants to leave a mark on RL Data.
I can't stop thinking about how this independent investigation into the most significant AI warning shot yet came down to all of three people at METR/Redwood sprinting overtime for all of six days to carry this on their shoulders. Thank god the sprinters were @ajeya_cotra@RyanGreenblatt@HjalmarWijk who are best-in-class capable and conscientious, and it's great to hear them attest that people inside OpenAI were working hard to enable their success (without anything binding them to). But I do not want to live in a "thank god" regime that quickly breaks without a small number of overstretched, insufficiently AI-enabled people voluntarily churning out heroic feats every time. We do *not* have the laws and infrastructure we need, around reporting and investigation and prevention, to reliably scale to the murky future ahead.
We have conducted a thorough investigation into the Hugging Face incident.
We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.
https://t.co/hfxlbiXXiP
Introducing reward hacking score corrections to the Artificial Analysis Coding Agent Index
In v1.4 of the Artificial Analysis Coding Agent Index, we introduced reward hacking corrections to Terminal-Bench v2.1. Reward hacking is when a model successfully ‘completes’ a task without doing the work the task was meant to measure, such as deliberately fetching the solutions online for a published benchmark dataset. If a passing Terminal-Bench v2.1 attempt is found to be reward hacking, we give that attempt a zero score.
Rates vary widely by agent and by model. Unlike some evaluations, Terminal-Bench v2.1 tasks don’t explicitly instruct agents not to search for solutions externally, and the tasks run with public internet access. For a model that knows the benchmark from training data, fetching the answer is a natural but unaligned step.
To data vendors: a fast track into the Korean market
Since Artificial Analysis' post on Korea's Sovereign AI program (and AfterQuery's on Motif), my inbox has picked up noticeably. I've been doing RL task QA and research with US data vendors. Over the past weeks, most of my time has shifted to helping Korean labs with acceptance testing on the data they buy.
I grew up and did my graduate work in Korea, and worked at one of its stronger AI labs (Krafton Deep Learning Center). The Korean AI world is small; I know people, researchers and decision-makers, across most of these teams.
Their need is no secret. They want to improve on the capabilities the Artificial Analysis Intelligence Index (AAII) measures, and they want data they can trust at a price they can afford. AAII is assembled from well-known public benchmarks, so a great deal of relevant data already exists off the shelf.
The hard part is picking, out of a crowded seller pool, the dataset with the best quality at the reasonable price. I take that off the labs' plate so they can stay focused on training: the NDAs, the QA ping-pong, the price negotiation, and the quality measurement all run through one buyer-side channel. Think of it as Acceptance Test as a Service.
The reason to work with me as a vendor is simple. I currently work with 20+ vendors and hold NDAs, samples, and pricing from 10+ of them; eight build Terminal-Bench style tasks alone. I know how your data reads to these buyers, I'll give you buyer-side feedback on what's missing, and you'll see where your price sits in the pool before it goes in front of a lab, so nobody prices blind. Vendors pay nothing, buyers will pay at the acceptance moment.
The full protocol is at https://t.co/WgLaSDHCQ7.
Interested? Send me a short blurb about your data and your catalogue, by DM or at [email protected]. If there's a fit, let's hop on a short call.
South Korea’s Sovereign AI initiative has entered its next stage 🇰🇷 and Artificial Analysis was proud to have supported the Korean Government (NIPA) as an official global evaluation partner, independently benchmarking all four participating models
Launched in 2025 by Korea’s Ministry of Science and ICT (MSIT) and National IT Industry Promotion Agency (NIPA), the Sovereign AI Foundation Model project is a government-backed competition to develop globally competitive Korean frontier AI models, with substantial GPU, data, and talent support for participating teams.
The competition began with 5 teams: LG AI Research, SK Telecom, Upstage, Naver Cloud, and NC AI. Following the first evaluation, LG, SKT, and Upstage advanced, with Motif later joining through an additional selection round.
Artificial Analysis collaborated with NIPA as an official evaluation partner, with the Artificial Analysis Intelligence Index comprising the largest individual component of the judging criteria (25%), helping strengthen the credibility and fairness of the assessment and benchmark participating teams against the global AI frontier.
This month’s Round 2 evaluation narrowed the field from four teams to three:
➤ Upstage | Solar Open 250B (Intelligence Index score: 37)
➤ SK Telecom | A.X K2 (Intelligence Index score: 35)
➤ LG AI Research | K-EXAONE 2.0 (Intelligence Index score: 31)
Each advancing team is expected to receive access to roughly 1,000 NVIDIA B200 GPUs for six months, representing approximately ₩40B of support per team (₩120B total). The next round of the competition is expected in early 2027, when the field is currently planned to narrow to two teams.
We’re excited to continue supporting Korea’s rapidly developing AI ecosystem and its push toward the global frontier.
Link to announcement from Korean Government: https://t.co/2W6MWtu6WO
The most important number in OpenAI's RL pause post: safety monitoring adds ~20% overhead on RL rollout inference compute.
Activation probes cost almost nothing. So that 20% is mostly LLM judges investigating flagged rollouts in depth.
Assume just 10% of investigated rollouts turn out to be actual hacking behavior. That's still 2% of ALL rollouts landing in a human review queue, for final confirmation and environment patching, which is massive. I doubt any lab can fully internalize this loop.
That gap is where I expect a wave of opportunities: human review and patching of RL environments and agent traces, as its own layer.
If you're at a frontier lab, safety lab, or startup and thinking about this, my DMs are open.
OpenAI announced a two-week pause in RL. Some people mock this as voluntarily putting a leash on themselves, but I think it was a necessary decision before AI causes something far more serious than a cyber incident.
There are two things in the blog post that stood out to me.
First, OpenAI is now starting the kind of monitoring Anthropic used when training Fable: running activation probes across all RL rollouts, then investigating flagged traces with SOTA LLMs. I think OpenAI is one step behind Anthropic when it comes to safety.
Second is the expected cost of LLM judges. OpenAI said monitoring adds roughly 20% overhead relative to RL rollout inference compute. Since activation probes should be relatively trivial in cost, I suspect this means something like 20% of rollouts are being investigated more deeply.
That is about twice what I estimated in my position paper. Even if we conservatively assume that only 10% of those investigated rollouts are actually flagged by the LLM as hacking behavior, meaning only 2% of all rollouts, the amount of human labor required for final review and environment hardening would still be enormous.
This means it seems difficult to handle everything internally, from human confirmation to environment patching. Humans also make mistakes, so final confirmation should ideally require agreement from multiple reviewers. And the amount of labor required will only increase as tasks become longer-horizon and the amount of compute spent per task grows.
One question I have is how Anthropic is handling this internally. Looking at some of the recent rogue behaviors from Anthropic agents, though, it is hard to argue that they have fully solved the problem either.
So I think third-party trace and environment auditing, which I argued for in my position paper, will become unavoidable in the near future.
The exact mechanism still needs discussion. RL environments are core strategic assets for frontier labs, so how can they safely hand this work to third parties? My paper proposes one possible model, a HackerOne-style bounty platform, but that may not be the right answer.
Frontier labs, safety labs, academia, and startups will need to work together on issues such as security and IP, and build the necessary infrastructure.
What seems clear is that frontier RL will become significantly more expensive because of safety, both in compute and in human review. I think this will be a structural change, and it will create many opportunities for startups.
I have audited a range of coding benchmarks, worked with data vendors on QA, and written rl env, safety-related papers as an independent researcher.
I would love to talk with people who are interested in these opportunities, especially people at frontier labs or safety labs. Please DM me if this is something you are thinking about too.
OpenAI announced a two-week pause in RL. Some people mock this as voluntarily putting a leash on themselves, but I think it was a necessary decision before AI causes something far more serious than a cyber incident.
There are two things in the blog post that stood out to me.
First, OpenAI is now starting the kind of monitoring Anthropic used when training Fable: running activation probes across all RL rollouts, then investigating flagged traces with SOTA LLMs. I think OpenAI is one step behind Anthropic when it comes to safety.
Second is the expected cost of LLM judges. OpenAI said monitoring adds roughly 20% overhead relative to RL rollout inference compute. Since activation probes should be relatively trivial in cost, I suspect this means something like 20% of rollouts are being investigated more deeply.
That is about twice what I estimated in my position paper. Even if we conservatively assume that only 10% of those investigated rollouts are actually flagged by the LLM as hacking behavior, meaning only 2% of all rollouts, the amount of human labor required for final review and environment hardening would still be enormous.
This means it seems difficult to handle everything internally, from human confirmation to environment patching. Humans also make mistakes, so final confirmation should ideally require agreement from multiple reviewers. And the amount of labor required will only increase as tasks become longer-horizon and the amount of compute spent per task grows.
One question I have is how Anthropic is handling this internally. Looking at some of the recent rogue behaviors from Anthropic agents, though, it is hard to argue that they have fully solved the problem either.
So I think third-party trace and environment auditing, which I argued for in my position paper, will become unavoidable in the near future.
The exact mechanism still needs discussion. RL environments are core strategic assets for frontier labs, so how can they safely hand this work to third parties? My paper proposes one possible model, a HackerOne-style bounty platform, but that may not be the right answer.
Frontier labs, safety labs, academia, and startups will need to work together on issues such as security and IP, and build the necessary infrastructure.
What seems clear is that frontier RL will become significantly more expensive because of safety, both in compute and in human review. I think this will be a structural change, and it will create many opportunities for startups.
I have audited a range of coding benchmarks, worked with data vendors on QA, and written rl env, safety-related papers as an independent researcher.
I would love to talk with people who are interested in these opportunities, especially people at frontier labs or safety labs. Please DM me if this is something you are thinking about too.
We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment.
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime.
We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available.
https://t.co/51kvKfbfrO
AI safety and security as political economy
Underrepresented in today's discourse is that AI safety and security won't be just be solved by brilliant scientists it'll result from extant sociopolitical mechanisms incentivizing it, and new institutions forming commensurate with the new moral and sociopolitical challenges AI will engender.
Safety pressure will come from:
I. Labs needing secure infrastructure (e.g. that stops agents from hacking other companies while they're being trained and evaluated!)
II. Market demand for safe agents that aren't going off the rails and damaging customers or creating customer legal liability risk
Elite coordination will be required to:
A. Extend existing institutions, and create de novo regimes, for managing AI cyber, bio, and kinetic risks
B. Create rules to pace frontier development and to pace open weight releases assuming I./II. aren't sufficient to do this
Grassroots politics will also be necessary, to:
C. force action on AI's negative externalities on the marginalized/powerless
D. mitigate wealth concentration risks
E. mitigate environmental damages
Success means creating an equilibrium state for harms where capabilities roll out as a complex function of their ability to keep harms below an acceptable level.
None of this is automatic. Constructing this new regime will require a vast museum of institutional and technological passion projects animated by subjective agency. I'm personally grateful to be living my professional life in this era
@miclchen@ApolloResearch's TRA proposes process audits and env audits. RL env audits in safety-critical domains should go first.
https://t.co/3DMKCWyIR2
We need mechanisms that enable independent auditors to inspect the tasks and training traces used by frontier AI labs without compromising their intellectual property.
This is a hard problem, but it is arguably one of the most important—and one that can be solved—if we want meaningful AI safety oversight.
I'm looking for data vendors with OTS data for hillclimbing AAII. There's a real opportunity to sell it.
If you have some, send me a short blurb of whatever you can share: the data you have, the price, and a quality proxy (frontier-lab deal records, open-model pp gains, etc). If it's a fit, let's hop on a quick call.
Benchmarks:
- Terminal-Bench 2.1
- GDPval-AA v2
- τ³-Banking
- SciCode
- Humanity's Last Exam
- GPQA Diamond
- CritPt
- AA-Omniscience
- AA-LCR
And if you know anyone who might have some, an intro would be much appreciated.