wrote up my thoughts on AI in 2026 and beyond - proactive agents, memory systems, on-going or new trends, and some harder questions about what happens when AI creates much more value than we do...
https://t.co/Fl5sjR4212
🚀 Hy4 preview is here.
770B, 49B active, 1M context.
Built for productivity.
Open source frontier.
Consistent affordable price.
Use it. Tell us what breaks.
More on Hy blog:https://t.co/rbl1IWRk3C
HuggingFace:https://t.co/mE9wevH5XR
Github:https://t.co/pyl9zckpoL
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length.
🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Try it now at https://t.co/GCdiMzk1Dl via Expert Mode / Instant Mode. API is updated & available today!
📄 Tech Report: https://t.co/drlDrxkYtp
🤗 Open Weights: https://t.co/T13Y8i7SDM
1/n
Two days ago, Anthropic cut off third-party harnesses from using Claude subscriptions — not surprising. Three days ago, MiMo launched its Token Plan — a design I spent real time on, and what I believe is a serious attempt at getting compute allocation and agent harness development right. Putting these two things together, some thoughts:
1. Claude Code's subscription is a beautifully designed system for balanced compute allocation. My guess — it doesn't make money, possibly bleeds it, unless their API margins are 10-20x, which I doubt. I can't rigorously calculate the losses from third-party harnesses plugging in, but I've looked at OpenClaw's context management up close — it's bad. Within a single user query, it fires off rounds of low-value tool calls as separate API requests, each carrying a long context window (often >100K tokens) — wasteful even with cache hits, and in extreme cases driving up cache miss rates for other queries. The actual request count per query ends up several times higher than Claude Code's own framework. Translated to API pricing, the real cost is probably tens of times the subscription price. That's not a gap — that's a crater.
2. Third-party harnesses like OpenClaw/OpenCode can still call Claude via API — they just can't ride on subscriptions anymore. Short term, these agent users will feel the pain, costs jumping easily tens of times. But that pressure is exactly what pushes these harnesses to improve context management, maximize prompt cache hit rates to reuse processed context, cut wasteful token burn. Pain eventually converts to engineering discipline.
3. I'd urge LLM companies not to blindly race to the bottom on pricing before figuring out how to price a coding plan without hemorrhaging money. Selling tokens dirt cheap while leaving the door wide open to third-party harnesses looks nice to users, but it's a trap — the same trap Anthropic just walked out of. The deeper problem: if users burn their attention on low-quality agent harnesses, highly unstable and slow inference services, and models downgraded to cut costs, only to find they still can't get anything done — that's not a healthy cycle for user experience or retention.
4. On MiMo Token Plan — it supports third-party harnesses, billed by token quota, same logic as Claude's newly launched extra usage packages. Because what we're going for is long-term stable delivery of high-quality models and services — not getting you to impulse-pay and then abandon ship.
The bigger picture: global compute capacity can't keep up with the token demand agents are creating. The real way forward isn't cheaper tokens — it's co-evolution. "More token-efficient agent harnesses" × "more powerful and efficient models." Anthropic's move, whether they intended it or not, is pushing the entire ecosystem — open source and closed source alike — in that direction. That's probably a good thing. The Agent era doesn't belong to whoever burns the most compute. It belongs to whoever uses it wisely.
Exclusive: Anthropic confirms it has begun testing its “most capable” model to date, after an accidental data leak revealed its existence
https://t.co/Mbc5b3Pb5Z
i made a tamagotchi that lives in your notch and reacts to your claude code sessions.
it cries when you yell at claude and gets happy when you praise it.
just added a page (https://t.co/ozkvjXCFpZ) for my favorite photos
these are cherry-picked ones and look great from my pov, but since i'm not a photographer, so maybe look bad lol
* and opus is sooooo good at design with tastes and consistency
from what this reveals, it's obvious to see how dishonest and unserious DoW is in series of actions
but despite of this, i think it's true that technology is much ahead than society (which includes legislation or smth), i expected this two yrs ago (https://t.co/zQuc8NAQyq.); however, we cannot expect laws to follow up quickly, so it has to be some organizations or individuals to hold the lines and examples at first.
again, it's really great to see how Dario and Anthropic are taking role here with principle; future would be bright with them (and one or two other players) taking the lead
i kinda have a feeling that even if we figured out how continual learning really works, it cannot be used in all those kind of "knowledge update" work.
it would be still limited to general and broad knowledge level updates (i.e. knowledge cutoff, learning a type of new task or smth); however, for individuals' tasks or usecases (your preferences, your codebase, your project context), better IF, better product and harness may just work better.
because i think weight updating per users has some reliability issue, apparently weights are not that easy to be interpreted, you cannot quantify how a model understands a person well, and if the model get something wrong, it may be difficult to modify that. but things like AGENTS/CLAUDE[.]md are transparent and you know what goes wrong, you could edit it easily.
We estimate that Claude Opus 4.6 has a 50%-time-horizon of around 14.5 hours (95% CI of 6 hrs to 98 hrs) on software tasks. While this is the highest point estimate we’ve reported, this measurement is extremely noisy because our current task suite is nearly saturated.
People should read the Claude Constitution. It does a pretty good job of laying out what Anthropic presumably really believes (and it is part of training). I’d think that a clear debate over things that are good or bad or missing there would be helpful. https://t.co/QU7aR8hxtD
> While many safety advocates warn about the dangers of humanizing chatbots, Askell argues we would do well to treat them with more empathy—not only because she thinks it’s possible for Claude to have real feelings, but also because how we interact with AI systems will shape what they become.
> A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”
The @DarioAmodei interview.
0:00:00 - What exactly are we scaling?
0:12:36 - Is diffusion cope?
0:29:42 - Is continual learning necessary?
0:46:20 - If AGI is imminent, why not buy more compute?
0:58:49 - How will AI labs actually make profit?
1:31:19 - Will regulations destroy the boons of AGI?
1:47:41 - Why can’t China and America both have a country of geniuses in a datacenter?
Look up Dwarkesh Podcast on Youtube, Spotify, Apple Podcasts, etc.
also, ML would be pretty much be partially (or even fully) automated within a year or two, as trend is pretty obvious through the METR's time horizon graph
seems like math would be one of the first few fields that would have fully automated research
and it kinda makes sense to me, because unlike physics, which needs understanding of real world, so may need some new architectures (or "world models"), math is something that only need rigorous text reasoning (even for geometry, you may need to transfer graphs to better-understandable text tokens or making model understand them better), so doing hard RL to current language models could push things even further.