Availability monitoring asks whether the job kept running. Correctness checks ask whether its computation can be trusted.
A service can return a result without crashing and still be wrong.
AI infrastructure needs both checks. Uptime alone cannot certify computational integrity.
A route record says what was documented. An observation says what was seen, where and when.
For the network carrying AI traffic, keep provenance, confidence and unresolved discrepancies with the state used to authorize changes.
A topology diagram is not physical verification.
Do not multiply gains from separate infrastructure models into a fleet-wide return.
The baselines may differ. Two fixes may recover the same lost hours. One change can alter the value of another.
A combined case needs one workload, one baseline and one accounting boundary.
Cost per token leaves out the work around the tokens.
For an inference service, count retries, tool execution and review against the number of tasks that meet the quality and latency target.
A lower token price does not by itself establish a lower cost per accepted task.
Separate prefill and decode pools create a handoff to operate: state transfer, capacity balance, admission and recovery.
Faster phase execution can be cancelled by a slow transfer or an overloaded pool.
Compare the complete serving path against a tuned shared fleet.
An inference service has at least two latency questions: how long until the first token, and how quickly the rest arrive.
Prefill and decode need separate measurements. Improving one phase can leave the user waiting on the other.
Capacity planning must meet both targets.
An action affecting another tenant's running job needs different authority from repairing your own slice.
Our governance design gives cross-hall actions explicit blast-radius limits and escalation rules.
A successful test does not erase those authority boundaries.
A fabric fault crosses more than one clock: transport retries, collective progress checks and the framework watchdog.
Changing one timeout does not change the others.
Our fault-fidelity work distinguishes these layers so tests target the mechanism they claim to exercise.
When satellite timing disappears, the site's clock becomes part of the availability budget.
Our reference AI-RAN model uses 6 hours of holdover for its OCXO and 36 for its rubidium clock within the chosen timing limit.
These are model inputs, not guarantees for every clock.
A backup path can preserve tenant traffic and still fail the radio workload.
In our AI-RAN model, bonded 5G and satellite paths cannot carry the radio class's roughly 100-microsecond fronthaul budget.
Availability must be evaluated separately for each service class.
A continuous 0.25 Gbps ingest stream produces about 81 TB in 30 days, before filtering or compression.
That volume appears in our edge-placement analysis.
A workload may need local processing because of backhaul cost or data controls even when latency permits a central site.
Our reference edge-fleet model strands about 87% of provisioned power when capacity is split across 30,000 towers in two-GPU units.
Power left over at many sites cannot become one larger allocation.
The result is scenario-specific. Aggregate megawatts hide allocation size.
A curtailment budget is a scarce resource too.
Across 24 synthetic price years, our scheduler saves more by targeting the highest-priced eligible hours than by curtailing a random set under the same budget.
The result is the ranking, not a promised saving on a real tariff.
Cooling has a water budget as well as a power budget.
In our direct-to-chip reference model, dry heat rejection trades lower water use for roughly 0.02 to 0.03 higher PUE.
Comparing systems on PUE alone hides a resource trade that the site still has to make.
Two halls can each have room for a job while their shared feeder cannot accept both starts at once.
Our admission model distinguishes a steady load that fits from a transition that exceeds the step limit.
Staggering starts can solve the second problem.
Rack density can eliminate a cooling option before PUE enters the comparison.
At 120 kW per rack, our reference ladder admits direct-to-chip and immersion cooling; its air and rear-door options fail the density gate.
Those are model envelopes, not universal vendor limits.
At 95% occupancy in our torus simulation, 38 of 40 failures reached the human-approval boundary because recovery could affect a neighbour.
At 25% occupancy, none did.
Recovery planning needs an approval-time budget as well as a spare-chip count.
In our placement experiment, dedicating a rack per tenant admitted 40,000 chips; open sharing admitted 75,392.
Isolation consumes placement options.
Some ring requirements also conflict with rack isolation. Reject incompatible requests before promising capacity.
In our 16 x 16 x 16 slice example, restoring a rectangle after one chip fails removes 256 chips at a corner, or 2,048 at the centre.
That is 6.2% versus 50% of the slice.
Location determines the modeled capacity loss. The other removed chips have not all failed.
In two induced-fault runs on one published optical testbed link, our analysis found FEC alarms about 88 and 39 seconds before outage.
That is evidence from one link, not a universal warning window. The question is whether the signal repeats on the paths a fleet depends on.