I’ve been testing GPT‑5.6 Sol in real agent workflows. What stands out: it follows long constraints, keeps using tools until the job is actually done, and pushes back on a bad premise instead of politely agreeing. The improvement feels practical, not cosmetic. Well played openai