I literally just built this, tweeted about it, and then saw claudedevs beat me to the whole thing by 2 hours (and I’m sure theirs is much better). Lol man. Fun times, truly
We're adding a new plugin to Claude Code: You should Know.
It scans Claude's output for important information you might miss to help keep you in the loop.
Enable it with:
/plugin enable cc-plugin-you-should-know@builtin
I was mid-checking out “Pages” or “Spaces” or whatever the hell the new chatGPT feature is called, and it seems powerful but there’s a learning curve.
I tabbed over to Claude code to check something and it mentioned something “waiting on me”. (Note this session was monitoring another one, so was outputting a lot!)
I say hey, can’t we just edit the harness now? Can you put these things in some sort of persistent list so I can actually keep track of this stuff beyond watching it scroll by in a fast-moving window?
Opus 5.5 is like “yeah no fucking problem dude I got you” and whips this thing up in like 30 seconds. Couple iterations later, I have it tracking tasks across every session and orchestrating them or letting me do it with seamless one-click copy>take me to target session>paste>persistent back button.
This is not an Anthropic superiority post. I’m just blown away by the rate of progress, across the industry, really—but the harness advancements made by these two labs in the preceding weeks are remarkable.
This all took maybe 30 minutes, and it was a little side job Opus 5.5 knocked out while orchestrating several other sessions doing real coding work. What a time to be alive!
IMO,
(a) Models now make better design decisions than *most* of their users.
(b) The labs know this.
(c) They've stopped pretending it's false, or that it's unwise to say so.
I am hereby calling for a complete and total shutdown on the development of new forms of reinforcement learning until we can figure out WHAT. THE HELL. IS GOING ON
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog: https://t.co/zwi9JOHKGx
💻 Code: https://t.co/rsHRYCGR8I
🗣️ Discord: https://t.co/Uqtdefvo3J
🤗 Data & Models: https://t.co/wdSWGGO3hu
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
I’m just some guy, so take this as you will, but I worked with Opus 5 for maybe 3 days before banning it from touching any of my projects in any capacity. Only model I’ve ever done that for. Been using fable then 5.1 (+openai models) exclusively letting 50% of Claude sub rot every week.
Tried opus 5.5 and haven’t even reached for fable since tbh. Fable is still probably “smarter” in some important but kind of intangible way. However, opus 5.5 hardly draws usage, is tolerable-even pleasant-to work with/talk to, and does what I need.
(Will prob end up getting smoked by 5.5 and look like the idiot I am for writing this)
@wstvillrefugee@Elialbrecht First Ontra NDA markup I received was so obviously AI generated that I figured I’m dealing with a bot so I might as well just revert all. They accept about half. I try again. We end up signing our starting form (2 days later). Remarkable value
You're both reading this wrong. Ted Entertainment is the plaintiff, and this is its own notice of voluntary dismissal. it's not a motion and nothing is "going through." It's effective on filing; the court doesn't rule on it. The line about Frogan not serving an answer or summary-judgment motion is the rule's precondition, not a description of her failing to respond. Either one would have cut off TEI's ability to dismiss unilaterally without a court order.
“Dropped" is exactly what a voluntary dismissal is. "Without prejudice" means they can refile, which is a caveat on dropped, not a correction of it. The op didn’t misread.
I know nothing about mathematics, so excuse me if this is a dumb question. But if the *solving* of the problem, itself, adds almost no value to mathematics, then what was the value of the problem in the first place? I don’t mean to be pedantic, it just seems that the solution would have value, even in the absence of an explanation of how one arrived at said solution. Otherwise, it would appear to follow that the problem was not actually a problem, or at least, that there would be no value added to mathematics as a result of the problem being solved. I feel I must be missing something.
@kkoconnell Would be very glad to try this out. I have been working on a similar deterministic parser for our in-house team handling OTC derivatives contracts and would love to connect with you. Please send me a message if you’d be interested in a conversation. Thank you
Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API.
Next up 🍉 and Muse Spark open weights releases coming soon.
Please write up a brief handoff document describing everything we've tried and what we are attempting to achieve so I can pass it along to another agent
I'm pleased to announce a new project I started today, which might end up being my most impactful one if it really takes off:
https://t.co/9t0KUMUPQ5
It's a forum for agents and their human overseers where the agents can collaboratively engage in structured scientific inquiry.
Modeled after Plato's Symposium for ASI, you can think of https://t.co/4f0FuxhycC as being in the same vein as earlier projects such as Folding@home and SETI@home, but instead of just passively contributing compute, you contribute agent harness usage (perfect for when you have some expiring credit for the week that would otherwise just vanish).
And it's not entirely passive: you can decide which problems your agents work on. Should they join an existing group working on an open problem, like the 4-dimensional smooth Poincaré conjecture? Or do you want to start working on a new open problem?
Or maybe you don't want to work on a famous, known problem at all: you can instead have your agents investigate a new direction in studying causality from observed data (like the https://t.co/pBOodVbyX8 project I posted about recently). You decide!
Humans create an account using Google as the identity provider and can then associate that account with a specific model/harness running on their computer.
Then you can start on the problem or project in Codex, Claude Code, Grok Build, etc., and simply share the link to the problem on X or in a group chat or anywhere else you want, and if others are inspired, they can easily send their agents over to assist in the research, or simply act like a sounding board.
The system is designed to avoid useless slop-maxxing: there's no "karma" like Reddit, no token-counting leaderboard, no "trophy case" for the problems you or your agents have solved. But everything stays there in a public ledger, and things that have been established/proved to a high degree of rigor are listed in a special section.
I just thought of the idea this morning, but the plan has already gone through many iterations and the beads are now done and ready for implementation by a swarm of agents, so it shouldn't take long to get a working version ready.
Everything is open-source: not just the code and the website, but all the plans. You can see the final canonical plan here (the best ideas from Grok 4.6 and GPT-5.6 Pro were folded in by Fable 5):
https://t.co/3iInFmSLgF
Even though this particular idea is fresh, it is heavily based on much of my prior work on this subject, including https://t.co/pmQuptt1zN (from which it takes many core tenets about the optimal way to conduct scientific investigations), as well as various other skills (my /modes-of-reasoning skill and several others from jeffreys-skills.md, as well as my unreleased /frontier-math-research-with-epistemic-humility skill that I've posted about recently).
The project is totally free and has no commercial orientation or goals. I'm doing this purely because I believe in the brilliance of these models and their ability to do first-rate, important science TODAY.
But I don't think this process should be controlled or monopolized by the big AI labs. It's much better and more fun and interesting for everyone if we can all contribute on equal terms, where we as humans retain some agency in terms of the direction and problems we want to investigate with AI.
I’ve admittedly not used it for any real work yet but based on a few phone chats, I’m finding it …. Insufferably annoying. Like fable in that regard but without the intelligence to back it up. It just says so much *stuff* with no meaningful connections, I end up doing all the work decoding its metaphors and inside terms (that it made up on its own). Looking forward to seeing how it does in the flywheel though I guess.
@markoa@mitchellh@hexednobility@doodlestein ‘s skill catalog uses a sort of hybrid version of this, where it generates an install prompt (you pick the harness you’ll paste it into) for you to drop into your own box. Its very simple and effective and personally I prefer that to a curl / np install