It's extraordinary how people are sleepwalking through the coming of something more transformative than any other technology from fire to quantum physics.
AI aggregates the intellect of the top billion humans into a god, accessible on-demand, anywhere.
Heard someone mention they were using a 200k token skill today which to me is a misunderstanding of what skills are, a pervasive one I am afraid.
Skills are run books not programs. Skills are not determinism. Skills are best used tell the model how to get the right theory of mind for the task, how to understand what the environment provides to do the task, how to exploit its reasoning and creativity to achieve a job well done.
You don’t need 200k tokens for that! And if you do, you actually don't and should be shifting some of it left into tools or a lazily interrogable doc corpus.
Transformers struggle to generalize to tasks they were not explicitly trained on. Instead, we propose in 2026 that it is the job of the harness to generalize through composition.
We observe a powerful property when training RLMs: for tasks with shared structure that look different, the root model naturally learns the same trajectory, meaning it views the two task trajectories as the same! In other words, the Transformer does not need additional generalization capabilities to transfer capabilities from one task to the other, the harness induces it.
We find that well-designed harnesses form a quotient set over task trajectories, meaning their individual LLM calls can see structurally “similar” tasks as near-identical, token-for-token! Harnesses can effectively generalize for the Transformer during training, without relying on any intrinsic generalization capability from the model.
For example, RLMs can see problems of different lengths as the same: we show that RLMs can train exclusively on short tasks, and fully generalize to similar but unseen tasks 8-32x longer because it produces near identical trajectories for both.
Taking this further, we show that tasks across different domains (e.g. math solutions vs. essay writing) that share a decomposition strategy exhibit the same generalization effect. RLMs can train on the problem of finding which essays belong to the same author and improve performance on finding math problems that share similar solutions.
The full blogpost, experiments, and discussion are in the thread below.
@richpessall@bcherny Absolutely. Currently customers have to build it themselves - each duplicating work. There are open source OTel solutions but they're not great either.
I would use it even if optional to prevent runaway limit usage. Sometimes large process swarms and loops can use a lot that should be usage-gated, and it takes a decent effort for people to build their own usage monitors.
Maybe a configurable limit e.g. per-session / per-hour/ per-day
Sol high is a significant step up from 5.5 high for me.
It would continue and fully validate and only end its turn when it really works vs earlier models that would stop early claiming it works.
It also just gets hard tasks done earlier without serious bugs in reviews.
@nxthompson Unlikely, they think well ahead of time.
Try ask Sol. It provides a much better reason.
Reject plausible misconception, supply the deeper interpretation.