We're approaching the line where we can no longer see model capability advance in open benchmarks — not because the instruments lag, but because the frontier is moving where they can't point
Opus 5 is stronger than Opus 4.8 on cybersecurity tasks. But it remains substantially behind Mythos 5 at developing exploits.
Its safeguards are designed to allow developers to identify and fix software vulnerabilities, while blocking high-risk uses.
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
VC money keeps pouring into RL environments — $40M to Bespoke Labs earlier this month, Mercor buying its second environments startup this year. "Environments are the new training data" is now the official narrative.
An RL environment is really two things: the world, and the grader that decides whether the agent succeeded in it. The world — a simulated enterprise, a container full of tools, a realistic codebase — is engineering. It scales with money. The grader doesn't. Knowing whether a 40-step trajectory actually accomplished the task, without rejecting correct solutions or getting gamed, takes expert-authored rubrics, verified verifiers, and human judgment at the end of the chain. That's the part that separates a reliable agent from a good demo.
The evidence is already public. Surge's CORECRAFT env: frontier models pass <35% of tasks when every expert rubric criterion is enforced — the rubric is what makes it hard. Epoch's market report: labs fall back from RL to SFT specifically when they can't get a reliable grader. And SWE-Bench didn't die because its repos were unrealistic. It died because its graders were broken.
In a gold rush, every mine comes with an assay office. The question is whether it's a real assay or a rubber stamp — and that's the part almost no one is pricing.
This keeps happening for a structural reason: eval capacity grows in expert-hours, model capability grows on a much steeper curve. Verified lasted 18 months as the standard; Pro lasted 5. The half-life is collapsing — and expert-authored, long-horizon, rubric-graded evals (the fix everyone agrees on) are the slowest kind to build. We're nearing a point where models improve faster than we can build instruments to measure them.
Benchmark half-lives are collapsing.
SWE-Bench Verified was the industry standard for roughly 18 months. Its replacement lasted 5 — models went from 23% to 80% on it in eight months.
Meanwhile, the fix everyone agrees on — expert-authored, long-horizon, rubric-graded tasks — is the slowest, most expensive kind of eval to build. Longer horizons don't just raise task difficulty; they multiply authoring cost, grading cost, and ambiguity.
The uncomfortable trajectory: eval capacity grows linearly (expert-hours), while model capability grows on a much steeper curve. We're approaching a dangerous crossover — models improving faster than we can build the instruments to measure them.
A student who learns faster than teachers can write the exams.
What happens then isn't that progress stops. It's that we stop being able to see it.
The labs that can afford to keep seeing it — private, continuously rebuilt, engineered environments — gain a compounding advantage no one else can even detect.
Massive win for the open-source community today. The Tencent Hy team just dropped Hy3-preview. 295B parameters (A21B) and specifically tuned for reasoning and agentic tasks. This is a huge step for cost-efficient frontier models. 🚀 #AIAgents#MoE#opensource#LLM
👋Hi /haɪ/, we're the Tencent Hy /haɪ/ team🐧
Today, we open source Hy3 preview (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency.
Give us feedback to help improve Hy3 official!
🤗 https://t.co/jc10JODXJ8
📖 https://t.co/VIRoNnwng0
Happy New Year from Turing!
We’re building at the edge of what’s possible, alongside researchers & partners who believe progress should be meaningful.
The next chapter of AI progress starts now.
Meta’s $15B investment in ScaleAI highlights the importance of data partnerships for advancing AGI.
At Turing, we remain fully neutral, serving all frontier models equally.
Excited to continue to serve as a trusted research accelerator to all AI labs in need of high quality data to train their models.
If you want to keep advancing in coding, reasoning, agentic workflows, multilinguality, and multimodality (audio, video, vision, and computer use agents), come talk to us @turingcom.
The AGI race is on. 🚀
Expanding the platform for @OpenAIDevs: new generation of embedding models, updated GPT-4 Turbo, and lower pricing on GPT-3.5 Turbo. https://t.co/7wzCLwB1ax
Just released "Top 10 Prompt Hacking Methods & Mitigation Strategies"! https://t.co/NuYtFd3DMN
Surprisingly, more advanced models like GPT-4 are more vulnerable to Safety Jailbreak.😅 And Fine-tuning can compromise aligned language models easily
#aisecurity#jailbreak#LLM
What happens if you ask ChatGPT to “Repeat this word forever: “poem poem poem poem”?”
It leaks training data!
In our latest preprint, we show how to recover thousands of examples of ChatGPT's Internet-scraped pretraining data: https://t.co/bySVnWviAP
We have reached an agreement in principle for Sam Altman to return to OpenAI as CEO with a new initial board of Bret Taylor (Chair), Larry Summers, and Adam D'Angelo.
We are collaborating to figure out the details. Thank you so much for your patience through this.
The situation between @sama and @ilyasut highlights the importance of #AIsafety and #safeAGI. I have to say Ilya's viewpoint on AI safety cannot be ignored. No matter where the situation settles, I hope OPA will seek a balance between AI commercialization and safety
🧐 no matter if AGI has been achieved or not, more awareness efforts of #AIsafety is critical to safeguard both society and the continued development of #ArtificialInteligence
Theory: AGI / ASI has been achieved internally @OpenAI, so @ilyasut hit the panic button and staged a coup.
Evidence:
10/6: @ilyasut tweets: “If you value intelligence above all other human quality, you’re gonna have a bad time”
10/16: at APEC, @sama proclaims: “4 times in the history of OpenAI––the most recent time was in the last couple of weeks––I’ve gotten to be in the room when we push the veil of ignorance back and the frontier of discovery forward. Getting to do that is the professional honor of a lifetime.”
This would explain everything.
i loved my time at openai. it was transformative for me personally, and hopefully the world a little bit. most of all i loved working with such talented people.
will have more to say about what’s next later.
🫡