$5.3M seed for @ExpanseCompute , led by @crane_vc with PXN Ventures. Expanse predicts exactly what an AI workload needs before it runs, so teams get as much as twice the work from the same hardware. Full story: https://t.co/xXDBxSt5Fo
Most cluster monitoring tools can’t tell you, within your allocations, whats being used properly. It tells you a node is allocated. It doesn’t tell you what’s happening inside the allocation. So a job that asked for 400GB and used 60GB reports as busy and you see ‘high utilisation’, but in reality there is 340GB sitting there doing nothing.
Across a cluster this reduces the number of jobs able to flow through your cluster. People then complain they dont have enough capacity, so you buy more, it takes 6-12 months and you still have the same problem.
We’ve measured this problem across various production clusters. Checking every job, what it asked for against what it went on to use. A lot of them reserved 2-3x what they used.
On one cluster the gap came to $8M of capacity nobody could see in a single month.
I’m talking about this and more at @stacresearch in New York in October, drop me a message if you’ll be attending :) - we will also be at the London event as well.
Meet the most promising companies from the latest @ycombinator batch 👀
One big takeaway:
AI agents need more than intelligence.
They need memory, identity, observability, compliance, insurance and power.
My latest @Forbes deep dive 👇
https://t.co/1IPkx2hqSi
Expanse (@ExpanseCompute) unlocks wasted GPU capacity. Submit jobs with the right resources. Optimize them to run faster. Debug failures in seconds. Cloud + on-prem HPC.
Congrats on the launch, @ismaeel_bashir_, @nkdem_b, @yafet_melake, and @erenzmendi03!
https://t.co/3lFcfFk7gS