@meggmcnulty When a training job goes slow, the first job is to blame or exonerate the network fabric — otherwise you waste hours chasing GPUs that were fine.
OOB + fabric visibility is what lets you tell “degraded path / silent stall” from a real compute problem.
Honest question for people running multi-node training:
When a job stalls, how long until someone notices — minutes, or the next morning’s loss report?
The expensive part of a GPU failure isn't the failure — it's the hour of recompute after it. Don't pay for compute twice. Clear rundown of how it works from @thenewstack.
If you're at RAISE July 8, this is the one to grab a seat for. Jordan moderating means no fluff and probably a few spicy questions 😄 — and Suresh brings the Clockwork take on keeping GPU fabrics fast and fault-tolerant.
There's a quiet truth behind every breakthrough AI company: its ceiling was set long before the first model was trained.
It was set by the infrastructure underneath. The compute it could secure, the capital that financed it, and the cloud architecture it ran on.
That's the trinity now shaping the AI era.
Compute, capital, and cloud have stopped being back-office concerns and have become strategic advantages. Get them right, and you can scale faster than competitors can react. Get them wrong, and even the best model may never reach its potential.
Which is why infrastructure has become destiny.
This July 8, five leaders at the center of that shift take the Master Stage to discuss how compute, capital, and cloud are redefining the economics of AI:
Stephanie Cohen (@Cloudflare) | Don Barnetson (@CredoSemi) | Suresh Vasudevan (@clockworkio) | Greg Matson (@solidigm) | @JeffDenworth (@VAST_Data)
Moderated by @JordanNanos (@SemiAnalysis_).
If compute, capital, and cloud set the ceiling, is your infrastructure strategy built to win, or simply to keep the lights on?
This is the final ticket release for RAISE Summit 2026.
Join the leaders shaping the future of AI infrastructure on the Master Stage.
🔗 Secure your ticket: https://t.co/eXTSRuwLs9
I woke up to a major Claude outage this morning. It's a stark reminder that as AI becomes an integral part of our daily lives, uptime can’t be an afterthought.
We need a robust resilience layer built across both the hardware layer and the distributed system layer—for both training and inferencing. AI is no longer a luxury; it’s core infrastructure.
Checkpointing feels free because it's "just writing to storage."
Then you measure it: the pause, the bandwidth contention, the GPUs sitting idle mid-write, multiplied across every checkpoint interval over a multi-week run.
"Free" is doing a lot of work in that sentence.
𝐌𝐨𝐬𝐭 𝐭𝐞𝐚𝐦𝐬 𝐭𝐡𝐢𝐧𝐤 𝐭𝐡𝐞𝐲 𝐡𝐚𝐯𝐞 𝐚 𝐆𝐏𝐔/𝐱𝐏𝐔 𝐬𝐡𝐨𝐫𝐭𝐚𝐠𝐞. 𝐀 𝐥𝐨𝐭 𝐨𝐟 𝐭𝐡𝐞𝐦 𝐚𝐜𝐭𝐮𝐚𝐥𝐥𝐲 𝐡𝐚𝐯𝐞 𝐚 𝐫𝐞𝐥𝐢𝐚𝐛𝐢𝐥𝐢𝐭𝐲 𝐚𝐧𝐝 𝐮𝐭𝐢𝐥𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐩𝐫𝐨𝐛𝐥𝐞𝐦.
At RAISE Summit, Booth#27A, the Clockwork team is running live demos showing how FleetIQ catches failing workloads in real time, keeps jobs running through grey failures and faults, and eliminates the checkpoint-restart cycle that quietly drains your training budget.
If you're scaling distributed training and want to see how this works — book a 1:1 before the slots fill up.
Space is limited.
👇
https://t.co/B7Rd6bQeqv
#AIInfrastructure #MLOps #FaultTolerance #GPUClusters
The most expensive thing in a GPU fleet isn't the GPUs that are running.
It's the ones that are powered, allocated, and idle, waiting on a straggler, a checkpoint, or a job that crashed an hour ago.
Utilization is a finance problem disguised as an infra problem.
Check it out! @Clockworkio co-founder Yilong Geng will be giving a lightning talk, "Beyond NTP," at OTel Community Day in Austin on June 20.
https://t.co/U81QDUAiz0
#OTelDay#OpenTelemetry