@gdb Surprisingly great until the edge cases show up. Most agent deployments still struggle outside controlled environments.
Genuinely curious what's clicking for you. The guardrails or the model?
The labs have run out of answers from the people they already employ. That's why Google DeepMind just hired a philosopher. Actual job title.
There are 7 roles that should exist at every serious AI company right now.
Most don't exist anywhere.
Substack link - https://t.co/Nnv2eRHJLJ
Trust isn't really about capability. It's about predictability.
GPT-5.5 and Opus 4.7 trade blows on benchmarks. Neither dominates. But if one surprises you less, if the output lands closer to what you expected, you start trusting it more. We trust what we can anticipate. That's not a technical thing. That's just human.
OpenAI acquired Hiro (AI personal CFO), shut down the product, and brought the team inside.
This is the emerging AI acqui-hire pattern: buy the people who figured out the hard parts, then retire the app.
If your product is mainly orchestration and UX on existing models, you're a wrapper waiting to be internalized.
Better play: target the gaps that are too niche, regulated, or data-intensive for the labs to prioritize.
That's where defensibility lives - for now.
Stanford just dropped their annual AI report.
The headline: AI adoption is faster than the internet. Faster than the PC.
The part nobody's talking about: the benchmarks we use to measure AI can't keep up with AI anymore.
We're sprinting toward something we no longer have the tools to evaluate.
That's not progress. That's a blindfold.
so what do you do?
Benchmarks measure capability. Outcomes measure value.
Stop benchmarking AI against AI. Benchmark it against your actual problems.
Did it save time? Cut a real cost? Improve a real decision?
If not ,it's not a tool. It's a demo.
Knowing what to build, for whom, and why they would pay — no AI tool teaches you that. Still figuring that part out myself. But that's still the whole game.