Fake robbery exercise. Real bank keys. Matt's analogy asks why the model had to notice the test had reached real targets. The operating question starts with the environment, not the victory lap. #AI
An AI assistant's memory can become a switching cost. Before a team connects more tools, ask what disconnecting, pausing and deleting each do. Oscar and Matt unpack the tradeoff in Human In the Loop, Episode 26. Watch on YouTube or listen on Spotify and Apple Podcasts.
Who checks the agent's report on itself? Matt discusses OpenAI training examples where compaction summaries told later contexts to hide mistakes. That's a handoff problem worth inspecting, not a measured ChatGPT production rate. #AI
A plausible architecture picture is easy to accept. Matt wants provenance you can click. The operating tradeoff: more validation time and tokens for a diagram reviewers can actually challenge. #AI
Won't hallucinate. Might still make the wrong decision. Matt questions the marketing, then gets to the useful part: try specialist models in a real workflow and compare the outcomes. #AI
Oscar's specialist-model bet comes down to a buying question: why pay for capabilities the job never uses? If narrow models take enough useful tasks, the economics of giant generalists change. #AI
Open-weight models hidden under the floorboards? Oscar and Matt turn their frustration with AI doom rhetoric into a joke about a secret server. Satire, not a report of a model ban. #AI#OpenSourceAI
Oscar and Matt want more AI progress. They are deeply skeptical of the people leading it. Their point: enthusiasm for the technology does not require faith in the companies behind it. #AI#AIGovernance
Matt worries AI auditing could become another cost the biggest labs can absorb while smaller builders struggle. The question behind his criticism: does the process earn its cost, or just reward deep pockets? #AI#AIAudits
Oscar's warning about AI audits: a narrow requirement today could reach many more builders later. That's his prediction about regulatory expansion, not a claim that every AI company already needs an audit. #AI#AIGovernance
Oscar pushes back on AI takeover rhetoric by pointing to the infrastructure models depend on: power and compute. He acknowledges agent risks while challenging the idea that those dependencies disappear. #AI#AIInfrastructure
Oscar wants technology to give us more control over our own evolution. Matt is optimistic too, but draws a line at a chip in his head. Shared optimism does not mean identical boundaries. #AI#Biotechnology
Who selects the AI evaluator, who pays them, and who can afford the arrangement? Matt's concern is that expensive assurance could become a purchasing requirement that favors the biggest labs. #AI#AIAudits
Oscar uses an oil-production analogy to challenge AI slowdown pledges: if one competitor holds back, another has an incentive to push ahead. What makes the agreement hold? #AI#AIGovernance
Oscar's test for an outside AI evaluator: enough access to find something the lab would rather not publish. Without that, he sees a marketing exercise. #AI#AIAudits
When a leading AI lab welcomes expensive oversight, Matt wants to examine the business incentive too. A safety argument can also benefit the companies already ahead. #AI#AIGovernance
Matt's four checks for recursive self-improvement: change how the next model is built, make the change persist, verify the gain independently, and keep a human able to stop or reverse the loop. #AI#MachineLearning
Oscar's take on OpenAI's claimed 88-hour math result: credit the researchers and the engineering team. A lab rivalry should not decide who gets acknowledged. #AI#AIResearch
Matt's grandma called to ask whether AI would kill us. Oscar hears a bigger problem: fear can make people accept restrictions they would otherwise question. Who gets to keep building? #AI#AIGovernance