I just stumbled upon the CMS of the AI era: I'm working on a personal project (https://t.co/vOWiheIvsw) and I had Codex setup the site with Astro. I had it create a system where each post is a markdown file, and codify the steps for publishing a new post in the AGENTS.md file. So now when I want to post, I just give Codex the text/image. It makes the post, updates the site, and redeploys. If it wasn't clear already, Wordpress is dead.
For 30 years I thought soap just killed germs, then I learned that it mostly just makes them less sticky. I just had Codex built me a visualization of this process: https://t.co/iOtdbvgWfe
(the things we do with 40% usage remaining and a weekly reset incoming...)
@alexgetmancom@thsottiaux@OpenAIDevs@ClaudeDevs I made a video about this happening to me in APRIL and I still haven't heard back after appealing, a thousand comments on that video from others in the same situation, it's probably not from using the harness but who knows... https://t.co/Da2ltg7GoU
@kelvinbuildss I’ve built a tool to do just that and yes I trust it b/c it’s built on a planning protocol that makes it safer than just hoping the model does the right thing.
@FlowHaa It’s great! it feels very magical to just tell codex to deploy it and it just works, even with databases and storage connected. I’d love to hear your thoughts if you try it out!
@alaphati_t I'm building Bahama (https://t.co/45V4AeEXrd). Cloud toolkit for agents. handles hosting, DBs, connection strings, everything you need to tell your agent "deploy it" and it’s done.
Open source, creates one standard that works with popular cloud tools, feels a bit like magic..
Today I’m launching Bahama, a cloud toolkit that helps coding agents deploy the apps you build.
It's open source, installs with one line, and works with cloud providers you already use (like @vercel and @neondatabase).
https://t.co/8q9HF5oNgb
Transformers struggle to generalize to tasks they were not explicitly trained on. Instead, we propose in 2026 that it is the job of the harness to generalize through composition.
We observe a powerful property when training RLMs: for tasks with shared structure that look different, the root model naturally learns the same trajectory, meaning it views the two task trajectories as the same! In other words, the Transformer does not need additional generalization capabilities to transfer capabilities from one task to the other, the harness induces it.
We find that well-designed harnesses form a quotient set over task trajectories, meaning their individual LLM calls can see structurally “similar” tasks as near-identical, token-for-token! Harnesses can effectively generalize for the Transformer during training, without relying on any intrinsic generalization capability from the model.
For example, RLMs can see problems of different lengths as the same: we show that RLMs can train exclusively on short tasks, and fully generalize to similar but unseen tasks 8-32x longer because it produces near identical trajectories for both.
Taking this further, we show that tasks across different domains (e.g. math solutions vs. essay writing) that share a decomposition strategy exhibit the same generalization effect. RLMs can train on the problem of finding which essays belong to the same author and improve performance on finding math problems that share similar solutions.
The full blogpost, experiments, and discussion are in the thread below.
In 2026, how you use the AI model matters more than its raw intelligence... but how do you use it better? With a harness. These guys took frontier models from 13% to 99% with a better harness and I made a video about it: https://t.co/tz4K6S10aK
Today, we’re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set.
[schema] makes an LLM think like a physicist. 🧵
Fable is the first model that can read my mind. As a model is doing UX work, I typically preview the changes and write out fixes to send on the next turn. Fable just fixes everything before I have a chance to ask.