@appliedcompute is the best actual Post-training/Finetuning company I came across in a long while. No relationship with them - just feel like people doing good job deserve a shoutout.
Kimi K3 full fine-tuning is live on AC2. Our memory optimizations reduced GPUs required per training replica by ~40%.
At nearly 3T parameters, Kimi forced us to rethink how we manage memory, communication, rollouts, and checkpoints. The result is a much more efficient path to training frontier-scale open models.
TIL some DCs have no spines https://t.co/I8xcKdNOgZ
you can then just use node internal NVSwitch fabric as the spine. pxn for the win https://t.co/1HOkCPi0IY
The biggest misconception about open source models is that it’s a pure cost play.
The pitch is usually: get 90% of the intelligence for 10% of the cost.
But with training, you can actually get better performance AND better economics on the things you care about.
It's a much more interesting trade and pushes the Pareto curve out rather than sliding along an existing one.
is no one using GPUDirect Storage (GDS)? almost no cloud has support for this out of the box but have ways of getting this working? what's the latest here?
https://t.co/cI8pAioeWd
Today, SkyPilot is out of stealth.
Building custom intelligence is now existential. We help frontier AI teams build intelligence faster by removing their biggest bottleneck: AI compute fragmentation.
Frontier teams like @appliedcompute, @AbridgeHQ, @hippocraticai, @hcompany_ai, and @nubank already run on SkyPilot, with 10x faster time-to-intelligence and double-digit increase in GPU utilization.
AI teams today get compute anywhere they can. They then firefight compute fragmentation across providers. Researchers burn time on workload setup. Infra gets paged when GPUs go down. Frontier teams build slowly even on the fastest compute.
@skypilot_org turns your fragmented compute into one AI supercomputer, so you run frontier workloads faster. Many users manage 10,000+ GPUs across providers with SkyPilot. GPU hours consumption has grown 6x in the last 6 months.
1/ We're launching SkyPilot Platform — the AI compute platform for frontier AI teams to manage large GPU fleets and accelerate building custom intelligence.
Optimized for fleet management, team governance, and frontier workloads — pretraining, post-training, multi-cluster serving, and sandboxes. SkyPilot open source users can switch to the platform with a server URL change.
2/ We've raised over $20M led by @Lux_Capital (@breeves08), with participation from @AmplifyPartners (@dauber, @lennypruss), @coatuemgmt, @FoundationCap (@ashugarg , @JayaGup10), @RaceCapital, @thehousefund, and top operators like @alighodsi (CEO, Databricks), @JeffDean (Chief Scientist, Google), @rauchg (CEO, Vercel), @amasad (CEO, Replit), @ClemDelangue (CEO, @huggingface) and more.
We're hiring across Engineering and GTM to deliver the platform for the next decade of AI.
Above all, I'm excited to be building with the incredible team we've assembled, along with my cofounders Zhanghao @Michaelvll1, Romil @bromil101, Scott, and Ion @istoica05.
If you firefight AI compute, let's build.
@DavidSHolz there's some amount of ray on top of slurm, slurm to get allocation and ray running as the distributed runtime
ray head node as a slurm job: https://t.co/lEOfhnVtBl
ray workers as another slurm job: https://t.co/xOM6rsOzm5
Introducing Windsurf 2.0.
Manage all your agents from one place and delegate work to the cloud with Devin - so your agents keep shipping even after you close your laptop.
1/ today we're releasing muse spark, the first model from MSL. nine months ago we rebuilt our ai stack from scratch. new infrastructure, new architecture, new data pipelines. muse spark is the result of that work, and now it powers meta ai. 🧵
AI Infra Meetup with SkyPilot and CoreWeave - what a night! Packed room, great conversations, and solid talks on scaling AI infra across K8s, Slurm, and neocloud.
🌟 Highlights: @lbz____ (@AIatMeta) shared how Meta unified 100k+ GPUs across Slurm clusters with SkyPilot. @Michaelvll1 + @kevin_mingtarja (@skypilot_org) walked through SkyPilot's new features and ran a live multi-cloud demo. @deok_filho (@CoreWeave) broke down SUNK and the SkyPilot x CoreWeave integration. Huge shoutout to all speakers!
Thanks to @CoreWeave@wandb for being amazing partners in making this happen. Excited to keep building this AI infra community. More events coming soon.
they should add a mode to claude where it never gives control back to me. `--allow-work-forever`
like yeah just go on that route and keep trying forever claude, if I notice something wrong I'll stop you
we wanna put claude in a humanoid bot, then prompt things like:
> go touch the border of the universe, then come back to tell us what's outside /never-stop
> what exists outside the simulation? /never-stop