AI clears two hours from your calendar.
What fills the gap: your lifeβor someone else's next request?
A faster task is a productivity gain. A shorter workday is a separate decision.
@DramaPreachr So the fix isn't "get better at saying yes" β it's deciding in advance what earns the reclaimed time, before the calendar fills itself in.
@DramaPreachr Right β and "asks first" is usually just "easiest to say yes to," not "most important." Nobody schedules time for the thing that actually needed those two hours back.
@MrSaaSBuilder The trickiest part isn't even the meetings β it's that "better branding" means nobody has to actually decide to add the work. It just arrives labeled as something you'd have said yes to anyway.
He asked ChatGPT for help with his taxes. Then a CPA checked the answer.
A number the chatbot treated as almost certainly correct still needed calculating and verifying. Other questionable figures hadn't been flagged.
That's the catch: how do you ask the right follow-up when you don't know what's missing?
Imagine checking your AI budget dashboard and finding a $465 charge you donβt recognize.
Rob Berger says that happened in his test. He checked the actual account: no matching transaction. ChatGPT then acknowledged a dashboard error.
He connected his accounts to make budgeting easier. He still had to check whether the transactions were real.
The $465 example starts around 7:06.
Google's own ad showed its AI answering a 9-year-old's question about the James Webb telescope with total confidence β and got the discovery wrong. NASA confirmed it. The market didn't wait for a correction: Alphabet lost $100 billion in hours. Confidence was never the same as being right.
@datachad Exactly β the real test isn't whether it says yes twice, it's whether the second check comes from something outside the model itself. Varghese "passed" a double-check run by the same model that invented it.
Asking the same chatbot to check its own answer isnβt an independent check.
In the 2023 Mata v. Avianca case, ChatGPT reassured a lawyer that a fabricated case was real-and cited legal databases to back it up.
A second confident answer still needs a real source.
Gemini 2.5 Pro is the pricier model - the one built to be the smart pick. Given one clear coding task, it skipped the single requirement that was actually asked for, and the free Flash model did it right. The tester's own conclusion: benchmarks don't tell you which model survives your actual workflow.
The chatbot gave an answer. The customer made plans. Then the company said the answer was wrong.
That's the part worth remembering when you ask a bot about money: you're the one who has to act on what it says.
@datachad Exactly. That's why AlphaEvolve was interesting β a faster matmul algorithm verifies itself. You don't need Google's word for it. Benchmarks need trust, artifacts don't.
An AI researcher once got a chatbot to sincerely apologize β in character β for releasing dinosaurs into Central Park. Ask it to be a squirrel instead, and it'll happily describe loving nuts in the first person.
It's not malfunctioning. It's not confused. It's doing exactly what it was built to do: predict a plausible next word, with zero interest in whether any of it is true.
That's the same engine giving you investment advice at 2am.
@GaryMarcus That distinction is exactly the hole in how we evaluate models. Benchmarks measure single-turn knowledge - unreliability shows up on step 100, with nobody watching. We test the wrong property and call the score progress.