found out with the @paperinstr team that hash based session routing for agent rl rollouts can deathspiral your samplers and kill your job's throughput.
for max sampling throughput, you generally want to operate close to the kv cache capacity. leaving too much headroom wastes capacity, but crossing the limit causes evicted rollout prefixes to be re-prefilled.
your job's max concurrent rollouts per replica is lower bounded by kv cache capacity:
max rollouts / replica >= kv cache capacity / kv bytes at max context length
in practice, each rollout's working set is usually much smaller than the max context. if average rollout residency is ~50% of max context, you can schedule maybe ~2x that (with headroom)
agent sessions with cumulative context are typically pinned to sampling replicas to maximize prefix cache hits. standard inference engines have no session awareness, so during environment interactions, a rollout's cached kvs become eligible for eviction. if the engine has a request queue from other concurrent rollouts and is under cache pressure, the engine will admit the new prefill and evict the cache of any "idle" rollouts.
hash based routing (e.g. vllm router's consistent_hash policy) is a popular load balancing choice for sampling replicas. hash based routing produces uniform session distribution in expectation, but does not guarantee perfectly balanced live session counts.
skew creates overloaded and underloaded replicas, and an overloaded replica can become "poisoned" with no escape valve. when rollouts on underloaded replicas finish, their replacement can hash to an already overloaded one.
cache pressure -> eviction -> re-prefill -> lower throughput -> bigger queue -> more cache pressure
a simple mitigation is using session-based least loaded assignment. new sessions are assigned to the replica with the least number of active sessions. if you conservatively cap in-flight rollouts to the per replica lower bound described above, least loaded placement entirely prevent eviction. if you intentionally oversaturate based on expected residency, even though eviction is possible, least loaded placement give the useful property that the replica under pressure becomes quarantined. new rollouts stop being scheduled on the replica under pressure until one of its existing sessions finishes, allowing the system to recover.
perfect least loaded session routing does require the router to know when a trajectory is complete so it can decrement its live session count. many generic inference routers don't expose a "release session" api because they are designed around requests, not sessions.
a throughputmaxxing rollout scheduler should dynamically change rollout concurrency, pause rollouts, or explicitly offload/evict cache by session based on kv cache pressure and environment interaction time. this is especially important for rl where trajectories can grow over time as the policy changes, or when tasks have substantially different length distributions.
Testing Stealth Model against GPT-6 Astra, Muse Spark 1.3, and Grok 4.7 one shotting EV drivetrain market study
Same harness, prompts, etc. Crazy results this is white-collar AGI.
Agents are using legacy Python packages to edit Word, Excel, and PowerPoint files, hitting API gaps, dropping into raw OOXML, and creating silent corruptions.
We fix that today with OSS Paper Office:
harnesses as another axis for scaling, and one way to do that is artificially creating new harnesses even tho it would never exist.
now everyone is massively synthesizing rl envs and i wonder how far we are away from doing the same thing in harnesses
a couple of interesting takeaways from the livestream:
1. coding tasks make up the bulk (~70%), followed by visual (13%) and general tasks (12%). chat RL accounts for only 3%. I wonder whether this mix reflects the distribution of real-world requests they see.
2. more than 20 harnesses, wow. i can’t even name that many. what exactly counts as a harness here?
Update every prior on post-training scaling laws:
They hit SoTA with ~45k RL tasks, nearly all synthetic, with $2.6m of Chinese compute (~$1m in US)
That's a 2023 era pre-seed! Human data is stifling the US frontier, and neolabs are prematurely scaling compute.
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
a couple of interesting takeaways from the livestream:
1. coding tasks make up the bulk (~70%), followed by visual (13%) and general tasks (12%). chat RL accounts for only 3%. I wonder whether this mix reflects the distribution of real-world requests they see.
2. more than 20 harnesses, wow. i can’t even name that many. what exactly counts as a harness here?
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju
age of research for post-training ended a few months ago: "the marginal return of engineering the data and environment pipeline substantially exceeds that of algorithmic novelty in post-training"
This episode is a learning‑focused podcast. We walk through the Kimi K3 technical report, hoping to appreciate “the beauty of technology” together. Yutao draws connections to more than ten related papers as he breaks down the report. Fair warning: he speaks really fast.
https://t.co/7D9O7cd5ff
Paper Instruments benchmarked sandbox providers for their RL rollouts, and E2B's tool execution came in up to 3x faster than every other sandbox provider they tested, reducing idle GPU time and training cost.
@paperinstr trains frontier models for knowledge work (consulting, finance, banking, law), running thousands of concurrent GRPO rollouts, each needing the solver and grader fully isolated from one another to prevent reward hacking.
Every rollout also needs to start from the exact same state in order to reflect the policy being trained accurately. It also needs to boot and execute fast, since the GPU sits idle waiting on the rollout to finish.
With @e2b, Paper Instruments isolates each rollout in a microVM, builds the initial state from a template, then snapshots the running sandbox, letting thousands of rollouts launch straight from the snapshot instead of booting cold, so they get consistency and speed while reducing training cost.
Read the full case study: https://t.co/iDc2LXKVN9
Had a great time designing DiligenceBench and its reference harnesses with the team @paperinstr x @thoughtfullab
It is a benchmark for evaluating AI agents on real public-equity research tasks, giving us a shared environment for testing models and harnesses against dense, task-specific rubrics.
https://t.co/boJbZGLbTX
Today, @paperinstr and @thoughtfullab are releasing DiligenceBench, an agent-first benchmark for long-form equity research. DiligenceBench is designed to be a simple, grounded playground for our rubric-based RL and harness optimization work.
Let's estimate the scale of the inkling final RL run?
An initial guess
O(1K GB300s)
~1 week
~$5M?
Would be a fun interview question for a post-training lab. Glad to see some info like this shared publicly.
Great work! @WeiboLLM
It’s absolutely impressive to see how far a 3B model can go at the frontier.
Also very happy to see that our write-up on train-inference mismatch helped :)
https://t.co/PWQ8TDcznk
wild results for the small frontier: "the 3B parameter budget is already sufficient to support highly compressed, long-horizon mathematical reasoning"
great to see Weibo using our mismatch stabilization work from last year :)