For folks not understanding where we are at with AI now - if the task is verifiable, it is now solved by AI.
Yes, earlier days you could argue that 'AI isn't that good' (e.g. see first attached chart here, where GPT-4o underperformed average accountants on tasks)
But now, it's a totally different story (only ~2 years later). The second chart is astounding. It took me a while to even understand it because it looks so odd. It compares manual accounting tasks to tasks complete with Opus 5.5. Basically Opus 5.5. solves all tasks almost instantly at a 100% accuracy rate, whereas a person takes a lot more time and gets a lot of things wrong. It makes the chart look entirely silly because the axes aren't even comparable.
So yeah, this is where we're at. If it is verifiable, it is solved. This is just how these model architectures work now.