@SargisChilingar Slowing down may be necessary. But we probably can’t assume everyone will.
If not everyone slows down, accountability has to advance alongside capability.
@Yoshua_Bengio Perhaps the deeper question is how civilization should prepare as increasingly autonomous agents begin to affect reality before we can understand or predict their behavior.
@TomDavidsonX The real-world iteration point is especially interesting to me. As AI learns by acting, checking and iterating in the real world, how do we verify what actually happened along the way?
@Yoshua_Bengio@TIME Safety by design is essential. As AI becomes more autonomous, we may also need assurance at execution — what it was authorized to do, what it actually did, and what changed as a result.
@ControlAI Whether ASI should be banned is one question. How any boundary around autonomous systems becomes enforceable in execution — and provable afterward — may be another.
@Alexranra@hilbertspaess Exactly. Rules can be torn down. Institutions can fail. That may be why accountability ultimately needs something deeper: the ability to prove what actually happened.
@WSJ Let’s slow down the Humanity vs AI debate. Civilization was never built on control alone. It also built systems of accountability for what actually happens.
We keep asking whether AI will replace people.
But the more interesting question may be how civilization absorbs AI once it becomes part of everyday life. New capabilities usually arrive before the rules around them. Then society builds the norms, boundaries and responsibilities that let them stay.
@finkd This separation is the interesting part.
If every consequential action must cross an independent Sentinel boundary, the next question may be whether we can prove what was authorized, what actually executed, and what changed.
@alexandr_wang The Sentinel architecture is the interesting part to me.
If an agent action matters enough to gate, it may also matter enough to prove — what was authorized, what actually executed, and what changed as a result.
@a_karvonen Self-explanation is valuable.
But explanation and proof are different boundaries.
A model may explain why it acted. The harder question is whether we can independently prove what it actually did.
10,000 agents, coordinated for 88 hours, under monitoring and isolation.
This is no longer just a model capability story.
It is an execution infrastructure story.
As autonomy scales, boundaries, controls, and proof have to scale with it.
This model represents a step-function improvement on many benchmarks, and its training is ongoing.
Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents.
Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
AI can now produce proofs.
That is extraordinary.
But as autonomous systems begin to do science, the next question may not only be whether the result is correct.
Can we prove how the proof was produced — by which agents, using what evidence, through which actions and validations?
Capability is scaling.
Accountability will have to scale with it.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Excellent breakdown.
The observe–act–verify loop is a real leap.
But as computer use moves into consequential workflows, the next boundary may be proof: verifying that an action worked isn’t yet proving what was authorized, what actually happened, and who is accountable when it matters.