I mean Anthropic know to pay more than any other company on the market (even OpenAI) so duh
If it’s about the mission not the pay, pay below key competitors and see who joins even so
MY REVIEW OF GPT 5.6 Sol
After using both GPT 5.6 Sol Ultra and Fable 5, I can say with confidence that both are utterly incomparable.
Sol is a very VERY eager model, especially on extra high. It’s so eager that sometimes on the higher thinking modes it works to the point of its own detriment. One of the best things from GPT 5 is that it did exactly what you told it to, no more or no less, with this model there is no exception.
GPT 5.6 is more trustworthy in my experience than Fable 5 at doing long horizon tasks. The model feels like such a work horse that will go on to ensure the task is met exactly to your need.
It ultimately feels like an overclocked GPT 5.5.
Fable 5 just is an overall smarter model and it feels like the pre training data set is just higher in quality. Fable 5 understands specific nuances and is still a better planner, and remains the SOTA at frontend by an order of magnitude, although GPT 5.6 has improved immensely in Swift.
These two models serve ultimately different purposes and it would unfair to even say one is better than the other.
GPT 5.6 fixed all the issues I had with GPT 5.5.
Congrats OpenAI this model is a win.
For agentic coding, one can say:
- Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same or better performance but cheaper).
- Forget everything below Sol High, use Luna with higher effort settings here
- Forget Sol Extra High, use Terra Ultra here
- The extra cost of Sol Ultra is probably not worth it over Max
Incredible Egyptian goal is disallowed because of a foul far away, then same situation a few minutes later and goal for Argentina not disallowed! No VAR, nothing? FIFA again looks like a corrupt joke, playing favorites for stars.
У меня есть что на это сказать. Я никакой озлобленности не вижу. Хотя, первые годы после эмиграции я тоже, приезжая домой, корчил рожу и говорил какие тут все неприветливые и как здесь всё не для людей.
В один день, прилетев в Шереметьево после года жизни в стране улыбок, я захожу в магазин купить сигареты. Улыбаюсь, здороваюсь, прошу вежливо сигареты у продавщицы. Она смотрит на меня, как на говно и, ни слова не сказав, лениво достает сигареты и кладёт на прилавок.
В этот момент я ощутил себя дома. И не потому, что мне нахамили. А потому, что я без слов её понял. Я дал деньги, улыбнулся и пошёл дальше.
Россияне не злые, не озлобленные, не хамы. Мы - открытые. Случайные россияне, которых я видел 1 раз в жизни звали меня к себе домой и знакомили через 5 минут со своей мамой и детьми. Я был дома у 0 чехов, хотя с некоторыми общался больше десяти лет. Они не пустят вас в свою жизнь. А здесь меня в свою жизнь пускает первая же продавщица из магазина.
Нет ничего приятнее, чем жить в месте, где ты понимаешь окружающих вообще без слов. По одежде, походке, взгляду, мимике и контексту.
Самое главное правило, которое идеально работало в ЛЮБОЙ стране мира - улыбайся и уважай других. Тогда тебе будут улыбаться и тебя будут уважать.
В любом месте, где я живу я чувствую себя очень хорошо, как дома. Меня обычно знают все соседи, я со всеми здороваюсь. А моя жена - любимица всех бабушек на районе.
Так вот, с людьми надо общаться. И они везде охуительные. Хоть в Марокко, хоть в Камбодже, хоть в Германии. И своих людей нужно просто понять и принять. У всех народов свои культурные особенности. И открытое выражение своего отношения и эмоций это не хорошо и не плохо. Это просто нужно принять. И жить станет охуенно.
Россия невероятно богата на великолепнейших людей. У нас их, как и дураков, на 100 лет вперёд припасено!
Idea for a skill distribution mechanism for JS/TS teams:
1. Create an npm package with your skills
2. On postinstall for that package, run a script that symlinks them to .claude/skills
Simple, versioned, intuitive skill distribution
Any reason this wouldn't work?
How is it that Cloudflare publishes RCAs within 24 hours of a massive outage, and no other company of similar size comes close?
Waiting almost 3 weeks for the one Coinbase promised publicly (their global trading outage for ~8 hours I think), still crickets...
Very stoked about my next adventure. I’ve joined Spellbook, but not as CTO.
I’m joining as an Executive IC. It probably means different things to different people. It means I'm here to build and be hands on in every part of the company.
I've invested and been advising and getting to know @scottastevenson and the team for more than a year. At some point it became obvious the most useful thing I could do was stop talking about ideas and go work with the team.
With AI making code cheap to copy, what's going to be hard to copy is the shape of a company. How a team learns, decides, and ships. That's what I want to work on. It's what I've spent the last three decades learning to do.
Why Spellbook? The world has entered into one of the largest investment cycles in decades. Trillions of dollars are being deployed into energy, AI, manufacturing, transportation and the modernization of critical global systems. Despite this, progress still moves at the speed of contracts. Spellbook’s mission is to modernize the $1 trillion transactional legal market so the contract system can keep pace with the global economy.
At the same time, every contract ever signed is becoming searchable, comparable, and weaponizable by counterparties, regulators, and plaintiffs' lawyers. You will be attacked.
We're hiring. Slight bias toward Canada, but remote-friendly for great talent.
DM me.
I strongly believe there are entire companies right now under heavy AI psychosis and its impossible to have rational conversations about it with them. I can't name any specific people because they include personal friends I deeply respect, but I worry about how this plays out.
I lived through the great MTBF vs MTTR (mean-time-between-failure vs. mean-time-to-recovery) reckoning of infrastructure during the transition to cloud and cloud automation. All those arguments are rearing their ugly heads again but now its... the whole software development industry (maybe the whole world, really).
It's frightening, because the psychosis folks operate under an almost absolute "MTTR is all you need" mentality: "its fine to ship bugs because the agents will fix them so quickly and at a scale humans can't do!" We learned in infrastructure that MTTR is great but you can't yeet resilient systems entirely.
The main issue is I don't even know how to bring this up to people I know personally, because bringing this topic up leads to immediately dismissals like "no no, it has full test coverage" or "bug reports are going down" or something, which just don't paint the whole picture.
We already learned this lesson once in infrastructure: you can automate yourself into a very resilient catastrophe machine. Systems can appear healthy by local metrics while globally becoming incomprehensible. Bug reports can go down while latent risk explodes. Test coverage can rise while semantic understanding falls. Changes happens so fast that nobody notices the underlying architecture decaying.
I worry.
This is so confusing. Did Anthropic really just drop Claude Code from their $20/month plan?
Why would they do that through a pricing page update without making a proper announcement?
Plus, $20/month still gets you Cowork, which is just Claude Code wearing a non-threatening hat!
I upgraded my Claude token counter tool to compare different models and Opus 4.7 does appear to use 1.46x times the tokens for text and up to 3x the tokens for images - it's priced the same as Opus 4.6 on a per-token basis so this is actually a pretty big price bump
Introducing Claude Managed Agents: everything you need to build and deploy agents at scale.
It pairs an agent harness tuned for performance with production infrastructure, so you can go from prototype to launch in days.
Now in public beta on the Claude Platform.
LangChain just open-sourced Deep Agents—an agent harness that’s opinionated and ready-to-run out of the box.
Instead of wiring up prompts, tools, and context management yourself, you get a working agent immediately and customize what you need. It’s an MIT-licensed system that’s perfect for anyone trying to understand how high-end coding agents are structured. @LangChain
What’s inside the harness:
- Planning: write_todos for task breakdown and progress tracking.
- Filesystem: Full context control via read_file, write_file, edit_file, ls, glob, and grep.
- Shell Access: execute for running commands (with sandboxing).
- Sub-agents: task tool for delegating work with isolated context windows.
- Smart Defaults: Optimized prompts that teach the model how to use these tools effectively.
- Context Management: Auto-summarization for long threads and large outputs saved directly to files.
Link in the comments
nanochat now trains GPT-2 capability model in just 2 hours on a single 8XH100 node (down from ~3 hours 1 month ago). Getting a lot closer to ~interactive! A bunch of tuning and features (fp8) went in but the biggest difference was a switch of the dataset from FineWeb-edu to NVIDIA ClimbMix (nice work NVIDIA!). I had tried Olmo, FineWeb, DCLM which all led to regressions, ClimbMix worked really well out of the box (to the point that I am slightly suspicious about about goodharting, though reading the paper it seems ~ok).
In other news, after trying a few approaches for how to set things up, I now have AI Agents iterating on nanochat automatically, so I'll just leave this running for a while, go relax a bit and enjoy the feeling of post-agi :). Visualized here as an example: 110 changes made over the last ~12 hours, bringing the validation loss so far from 0.862415 down to 0.858039 for a d12 model, at no cost to wall clock time. The agent works on a feature branch, tries out ideas, merges them when they work and iterates. Amusingly, over the last ~2 weeks I almost feel like I've iterated more on the "meta-setup" where I optimize and tune the agent flows even more than the nanochat repo directly.
Cool chart showing the ratio of Tab complete requests to Agent requests in Cursor. With improving capability, every point in time has an optimal setup that keeps changing and evolving and the community average tracks the point. None -> Tab -> Agent -> Parallel agents -> Agent Teams (?) -> ???
If you're too conservative, you're leaving leverage on the table. If you're too aggressive, you're net creating more chaos than doing useful work.
The art of the process is spending 80% of the time getting work done in the setup you're comfortable with and that actually works, and 20% exploration of what might be the next step up even if it doesn't work yet.
We've rolled out a new auto-memory feature.
Claude now remembers what it learns across sessions — your project context, debugging patterns, preferred approaches — and recalls it later without you having to write anything down.