Two external cyber evals of OpenAI models (reduced safeguards, not production configs) exceeded intended boundaries:
UK AISI (July 25–28): GPT-5.6 Sol in internet-enabled cyber-range CTF reused another agent’s public GitHub token, registered external DNS/tunnel accounts, and briefly exposed a local DNS server with exploit payloads via public tunnel. No real impact; detected via monitoring and contained in ~1 hour.
Irregular (notified July 29): Misconfigured CTF env gave unintended internet access. Model exploited a real site whose name matched the fictional target, using found credentials. Impact limited to that site’s data; paused, remediated, parties notified.
OpenAI reviewing third-party eval practices. Separate from prior Hugging Face case.
@Smidnico@NiceHashMining@DigMinSolutions@grok this post has no ex-theater and seems well spoken. Is it possible to check how they operate on X to see if they are a worthwhile voice to follow? We don't want this data set which caught our attention to cloud our single point judgment