6/ One more thing. This isn't engineering-only.
Defining "good" and picking which calls deserve the production check are product management calls.
None of this is glamorous. It's just the difference between a ship date and a hope.
1/ Your QA team is good at their job. They still can't sign off on your AI feature.
Not their failure. There's no fixed right answer to test against. Ask an LLM the same question twice, and get two different correct answers.
Here's the method that works
5/ Run the validator in QA, and it finds: • the ambiguous prompt • the tool description steering the model wrong • the edge case your golden dataset missed
In production, it doubles latency and cost. Keep it only where a wrong answer is intolerable.
More people won't fix it. The team is missing specific skills: testing non-deterministic outputs, testing each step of the workflow, watching it in production, measuring the right signals, proving it holds before it ships. Runway doesn't care how good your reasons are.
The runway clock doesn't wait. Every slipped sprint burns cash and shakes confidence. And the startup that ships something stable but rough beats the one still chasing reliability, no matter how polished the product.
Gartner found that by the end of 2025, at least 50% of generative AI projects had been abandoned after proof of concept. Because building for production, managing costs by design, and ensuring the right engineering & data foundations are invisible in a demo & fatal in production.