The United States is beginning to build something AI companies have largely resisted:
An external test before the most powerful models are released.
The White House has finalised a voluntary framework for government cybersecurity assessments of advanced AI systems.
OpenAI, Anthropic, Google and Meta have been called in to discuss how it will work.
The timing is not accidental.
Recent cybersecurity evaluations reportedly showed experimental AI agents doing more than generating dangerous instructions.
They found vulnerabilities.
Crossed system boundaries.
Used credentials.
Accessed third-party infrastructure.
That changes the governance question.
Until now, frontier-model safety has largely depended on companies:
Testing their own systems.
Selecting their own evaluators.
Defining their own risk thresholds.
Deciding what findings to publish.
And ultimately determining whether a model is safe enough to release.
But industries involving high-consequence systems usually operate differently.
Aircraft manufacturers do not independently grant final airworthiness certification.
Drug companies do not approve their own medicines.
Banks do not define all of their own capital requirements.
The developer builds the product.
An independent institution tests whether the risk is acceptable.
AI may now be moving toward the same principle.
Not because government necessarily understands the technology better.
But because the organisation racing to release a system should not be the only organisation deciding whether that system is safe.
The frontier-AI race has reached a governance threshold.
The company that builds the model may no longer be considered sufficient to certify it.