Most engineering teams treat observability as a post launch. Build first, instrument later mindset.
Then production breaks and you're debugging with logs that weren't designed to tell you anything. Observability isn't a feature you add. It's a discipline you either have from day one or you pay for later.
@MWeckbecker Super interesting! I'm considering how this might play out without the intentional seed. Could a hallucination seed Agent0. Thought this was interesting "However, we do not observe the effect when the teacher and student have different base models." Thank you!
AI loves reading a 700 line markdown. People don't. Go figure. Creativity is still still best person to person. A huge markdown in a Slack chat is an antipattern, not collaboration. Thoughtful synthesis is still king in communication.
@SheetalJaitly Absolutely an organizational change problem. I'll boil it down one step further. Fear. More than just the fear of change. The fear of the unknown, or worse the dread of irrelevance. The new abundance is anything but clear at this point.
AI governance gets added as a checkbox to pass change boards or for regulators.
By then, you're retrofitting audits onto systems that were never designed for accountability. Bolted-on governance costs 10x what built-in governance costs — and still doesn't work as well.
Building governance in is hard in the short-run, but pays off in the long-run.
The cloud vs. on-prem debate for AI workloads isn't about GPU price per hour.
It's about workload predictability. Spiky inference: cloud makes sense. Steady-state training: reservations or on-prem. Treating both the same is how you overpay for both — and underperform on each.
I've spent weeks overthinking whether to talk openly about what I'm building. Then realized: nobody cares about your idea until you've proven you can execute it. The risk of staying quiet is higher than the risk of someone stealing a concept. Execution is the moat.
The conversation is shifting from "can AI write code" to "who owns what the code does after it ships." That's the right question. Throughput without accountability is just faster accumulation of things nobody fully understands.
AI coding tools get you to 80% in a day. Then you spend months in the last mile — reverse-engineering what was built, handling the edge cases it missed, debugging production incidents with no context for why the code exists. That's not a model problem. That's a specification problem.
GPU prices make headlines. Price doesn't equal value. Scheduling overhead, idle cycles between jobs, wasted compute capacity. Getting value out of GPU investment is where the work starts.
The best engineering teams I've worked with had headted arguements about architecture and design. Real disagreement about tradeoffs, constraints, and what "done" actually means. Then, agreement or not, they built. Win or loss, they were wiser and more cohesive in the end.
Demis Hassabis just defined the real test for AGI. It’s more brutal than anyone expected.
Train AI on all human knowledge. Cut it off at 1911. See if it independently discovers general relativity like Einstein did in 1915.
If it can, we have AGI. If not, we’re still building pattern matchers.
Hassabis: “My definition of AGI has never changed. A system that can exhibit all the cognitive capabilities that humans can.”
Not bar exams. Not coding competitions. All cognitive capabilities.
Hassabis: “The brain is the only existence proof we have, maybe in the universe, of a general intelligence.”
That’s why DeepMind studies neuroscience. Not for inspiration. For data. The human brain is the only confirmed evidence that general intelligence is physically possible.
If you want to build it, you study the only example that exists.
Hassabis: “True creativity, continual learning, long-term planning. They’re not good at those things.”
Current systems are impressive and broken simultaneously.
Hassabis: “They can get gold medals in international math olympiad questions, but they can still fall over on relatively simple math problems if you pose it in a certain way.”
Jagged intelligence. Brilliant in narrow domains. Incompetent when approached differently.
That inconsistency is the tell. A true general intelligence doesn’t spike in one direction and collapse in another.
The Einstein test cuts through all of it. No benchmarks. No leaderboards. No carefully curated evals.
Just a model, a knowledge cutoff, and the question of whether it can do what one human did alone in 1915.
Hassabis: “Training an AI system with a knowledge cutoff of 1911 and seeing if it could come up with general relativity like Einstein did in 1915. That’s the true test of whether we have a full AGI system.”
Current models can’t. They remix brilliantly. They don’t generate paradigm-shifting theories from first principles.
Hassabis: “I think we’re still a few years away from that.”
A few years. Not decades.
The system that can be Einstein once can be Einstein a thousand times simultaneously across every domain.
That’s not AGI anymore. That’s the beginning of something we don’t have words for yet.
When that test gets passed, we won’t need a press release to know what happened.
The danger in thought leadership is that it's presented or taken as "the answer". Answers have a radically shorter half-life these days. I like question leadership. The curiosity to know when to ask new questions and the humility openly seek answers.