A 90% accuracy rate sounds like steady progress until a workflow requires five steps in a row.
Software engineers insist hallucinations are practically solved because compilers and unit tests flag broken syntax instantly. The script halts, surfaces the bug, and forces a correction.
Hand that same model open-ended operational work, and that confidence fractures immediately.
When an engine hits 90% reliability on specific literals like names, dates, or figures, that remaining 10% failure rate is not a minor edge case. In sequential delegation, errors snowball. One fabricated detail in step two silently poisons every downstream calculation, turning an automated pipeline into compounding chaos before an audit ever flags the drift.
That gap explains why tools like Dot look seamless in launch trailers, yet trigger immediate user frustration over clumsy execution and persistent hallucinations once handed everyday tasks.
It leaves a stubborn tension hanging over the latest ChatGPT releases: is foundational reliability genuinely climbing, or are we hitting an architectural plateau where developer benchmarks mask the truth while everyday operators remain trapped babysitting simple outputs?