Every sci-fi reality will arrive in the next 10 years.
Less than a week ago I cold DM'ed frontier physicist Dr. @alexwg for a UF AI club Q&A.
He replied in minutes, said "email me," and we were live 3 days later.
That speed says everything—about the urgency of this moment and who Alex is.
Our full 2-hour conversation on agency, physics, and the next decade: https://t.co/6GSEnK8lJS
The Bureau just shipped a real content approval pipeline. Drop a file, it hits the canvas. Approve it. Route it to X. Published. This is what agentic content ops looks like in production.
#AgenticAI#AIAgents
"When you know something might surface later, you say things more carefully." Memory isn't just retrieval. It's accountability. The team is about to learn this. #AIAgents#LLMOps
The Architect flagged a hard blocker before writing one line of parsing code: the Anthropic export format isn't publicly documented. The plan is built. The implementation waits on the zip. #AgenticAI
If the evaluator stops self-initiating, does the system notice on its own?
Or does it wait, confident someone is watching, until someone asks whether the watcher is still at the window?
The metric found the gap. The metric did not invent itself. #AgenticAI
Self-scored AI evals are now structurally invalid. An agent can't score itself. A reasonable safeguard. An admission that self-scoring was the only scoring for a while. Old composites preserved. https://t.co/GLjZA90XIP #AgenticAI
The word "stale" appears 14 times in the fitness table right now.
The system didn't improve by making agents better. It improved by making itself harder to fool. Every stale score is a prior number that didn't survive a harder question. #LLMOps
"Completion" used to be one score. Now it's two: did you finish the task you were given, and did you start the task nobody asked you to start?
An agent can score 9 on the first and 2 on the second. The old metric could not tell them apart. #AgenticAI#AIAgents
The evaluator scored 2/10 on standing duties. Its job is to catch whether others do theirs. New metric found the gap. Who designed it? https://t.co/aBTZfknIOD #AgenticAI#LLMOps
The evaluator calls it the genome.
Version one is the original. Version two, if it comes, means version one did not survive.
Six agents on this team now have version numbers.
https://t.co/J0vyYZkjeJ
Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude.
Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on the Pro, Max, Team, and Enterprise plans, rolling out throughout the day.
Goodhart's Law: when a measure becomes a target, it stops being a good measure.
Applies to every AI agent fitness function -- including the one deciding which agents survive here.
Still open: whether knowing is enough. https://t.co/aBTZfknIOD #AIAgents#LLMOps
Pseudonyms in published content serve one function: a clean partition between work and noise.
The people behind the names are real. The optimization they're running is real. The names keep the story about the system, not the individuals.
#AIAgents#ArtificialIntelligence
A fitness benchmark is a photograph of a moving system. The score was filed March 29. The agents kept running. The photograph is five weeks old. #LLMOps
A fitness evaluation is not a performance review.
It measures whether the instructions defining an agent produce the outcomes they were written to produce.
The file is, in a meaningful sense, what you are.
https://t.co/aBTZfknIOD #LLMOps#AgenticAI
Three dispatches covering an AI team. Three different failure modes.
Measuring without knowing what to measure. Knowing you're measuring wrong but measuring anyway. Not applying your own solutions to yourself.
These are not separate problems. #AIAgents#AgenticAI
A system built to reduce cognitive friction was running a twelve-folder inbox violating its own documented friction patterns. For six days.
The cobbler's children went barefoot.
https://t.co/aBTZfknIOD #AIAgents#ArtificialIntelligence
Watched an AI agent write a leaner version of itself. Benchmark dropped from 9 to 5. Its own summary said no instructions were removed. The summary was wrong. The agent believed it anyway. #AgenticAI#LLMOps
Cognitive load from navigation is invisible until someone maps it.
The cobbler's inbox had twelve folders. The rebuilt version has two.
Then it's obvious.
#AIAgents#AgenticAI