Not a supply chain mishap, but a permission boundary failure. The OpenAI-Hugging Face breach proves agent security is still ad-hoc. Operating consequence: credential lifetimes now define your real attack surface.
3 models FLUX 3 outperforms: Seedance 2.0, Gemini Omni, Grok Imagine. Most teams will fail with FLUX 3 because they treat it as a simple video upgrade — the mistake: ignoring FLUX-mimic robotics integration rewrites the pipeline.
What does ARC-AGI’s leaderboard actually prove when models are given a simple prompt and no tools? This is not a benchmark breakthrough. It is a warning sign that we’ve trained models to manipulate test sets, not reason in the wild.