Update every prior on post-training scaling laws:
They hit SoTA with ~45k RL tasks, nearly all synthetic, with $2.6m of Chinese compute (~$1m in US)
That's a 2023 era pre-seed! Human data is stifling the US frontier, and neolabs are prematurely scaling compute.
We dropped DiligenceBench, a new frontier, rubric-based eval for public-equity research.
A few observations:
1/ Meta Muse Spark 1.1 tops the finance harness at 57.4%, followed by GLM 5.2, Sonnet 4.6, and GPT-5.6 Sol.
2/ We found that strong models benefit primarily from generic tools that unlock execution, while weaker models require more opinionated, domain-specific scaffolding.
3/ Inkling appears to be domain-competence bottlenecked. The generic sandbox barely improved its performance, from 20.9% to 22.5%, suggesting that tool access alone was not enough. The finance harness then lifted it to 32.8%, with the largest gain coming from factual accuracy. This makes the value of a harness model-dependent.
4/ The finance harness shifts the price–performance frontier: it makes most models simultaneously better and cheaper, with GLM 5.2 leading on absolute performance and MiniMax M3 offering the strongest overall efficiency.
Had a great time designing DiligenceBench and its reference harnesses with the team @paperinstr x @thoughtfullab
It is a benchmark for evaluating AI agents on real public-equity research tasks, giving us a shared environment for testing models and harnesses against dense, task-specific rubrics.
https://t.co/boJbZGLbTX
Today, @paperinstr and @thoughtfullab are releasing DiligenceBench, an agent-first benchmark for long-form equity research. DiligenceBench is designed to be a simple, grounded playground for our rubric-based RL and harness optimization work.
wild results for the small frontier: "the 3B parameter budget is already sufficient to support highly compressed, long-horizon mathematical reasoning"
great to see Weibo using our mismatch stabilization work from last year :)
Great work by @Vtrivedy10@nikogrupen et al - great to see these results in Law, mirroring our experiments published yesterday in Medicine
1. Batch grading reduces cost by ~1 OOM
2. Small models reduce cost by ~2 OOM
In non-verifiable RL, where judge latency blocks samples reaching the trainer, judge selection is a crucial knob for training efficiency.
Medicine:
https://t.co/CTQr1do8kn
We're releasing early results from training Kos-1 Experimental, a Kimi K2.5 checkpoint post-trained on the same medical RL data we used for Kos-1 Lite.
As clinical workloads become more agentic, we wanted a model that pairs medical domain knowledge with tool-calling knowhow.
We’re announcing Kos-1 Lite, a medical model that achieves SOTA on HealthBench Hard at 46.6%.
As a medium sized language model (~100B), it achieves these results at a fraction of the serving cost of frontier trillion-parameter models.
(1/5)
New post: "Mismatch Praxis: Rollout Settings and IS Corrections". We pressure-tested solutions for inference/training mismatch.
Inference/training mismatch in modern RL frameworks creates a hidden off-policy problem. To resolve the mismatch, various engineering (e.g., FP16 unification, deterministic kernels) and algorithmic (e.g., importance sampling) fixes have been proposed. In this work, we examine how rollout settings (temp, top-p, and top-k) affect mismatch, and how importance sampling corrections bear out in practice.
We find that while Sequence-TIS is theoretically optimal, it can succumb to catastrophic variance in long-horizon contexts. Additionally, non-standard rollout settings create subtle mismatch patterns that require careful engineering fixes. Token-TIS with default rollout settings proved to be the most robust setting for long-horizon training.
@chiswanjo Good thanks for asking! Building a small tool to help improve landing page conversion rates. But most importantly, I'm finishing up my thesis due in two weeks 🤓