📢
#NSERC launches the Geoffrey Hinton Prizes!
These prizes recognize Discovery Grant applicants leveraging #AI in novel, innovative and impactful ways.
Learn more ▶️ https://t.co/i511DM1noq
#NSERC_Prizes
People are talking to AI models as if they were romantic partners. That's not sci-fi, it's research data.
I asked Kai-Fu Lee if that counts as real love — or just a convincing imitation. His answer mixes scepticism with unexpected empathy.
Full episode on YouTube 🎙️
New uses require new interfaces
These are the three things I have been trying out to control Codex. Given how useful voice mode is, the Teenage Engineering Ting walky-talky has been a surprising favorite. The Codex micro is beautiful & fun, but too little info vs the Stream Deck
Today, SkyPilot is out of stealth.
Building custom intelligence is now existential. We help frontier AI teams build intelligence faster by removing their biggest bottleneck: AI compute fragmentation.
Frontier teams like @appliedcompute, @AbridgeHQ, @hippocraticai, @hcompany_ai, and @nubank already run on SkyPilot, with 10x faster time-to-intelligence and double-digit increase in GPU utilization.
AI teams today get compute anywhere they can. They then firefight compute fragmentation across providers. Researchers burn time on workload setup. Infra gets paged when GPUs go down. Frontier teams build slowly even on the fastest compute.
@skypilot_org turns your fragmented compute into one AI supercomputer, so you run frontier workloads faster. Many users manage 10,000+ GPUs across providers with SkyPilot. GPU hours consumption has grown 6x in the last 6 months.
1/ We're launching SkyPilot Platform — the AI compute platform for frontier AI teams to manage large GPU fleets and accelerate building custom intelligence.
Optimized for fleet management, team governance, and frontier workloads — pretraining, post-training, multi-cluster serving, and sandboxes. SkyPilot open source users can switch to the platform with a server URL change.
2/ We've raised over $20M led by @Lux_Capital (@breeves08), with participation from @AmplifyPartners (@dauber, @lennypruss), @coatuemgmt, @FoundationCap (@ashugarg , @JayaGup10), @RaceCapital, @thehousefund, and top operators like @alighodsi (CEO, Databricks), @JeffDean (Chief Scientist, Google), @rauchg (CEO, Vercel), @amasad (CEO, Replit), @ClemDelangue (CEO, @huggingface) and more.
We're hiring across Engineering and GTM to deliver the platform for the next decade of AI.
Above all, I'm excited to be building with the incredible team we've assembled, along with my cofounders Zhanghao @Michaelvll1, Romil @bromil101, Scott, and Ion @istoica05.
If you firefight AI compute, let's build.
Grok Voice Think Fast 2.0 is built for voice agents in the real world. It hears clearly in noisy conditions, reasons through complex workflows, and sounds more natural in conversation.
Introducing Runway Media Router.
The first preference-optimized router for generative media. Instead of hand-picking a model for every request, you define what "best" means once, for cost, quality, or latency, and the router selects the right video, image, or audio model automatically.
Live now in Runway Dev.
Numbat runs across our systems and sends audit events, findings, and security alerts to Computer.
Computer continuously analyzes this telemetry, escalates suspicious behavior, and proposes new on-device detections for Numbat, creating a continuously improving detection flywheel.
New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taught by @zhennydez, @duerr_seb, and @MilksandMatcha.
When a model generates text, much of the time is spent moving its weights out of memory and into the compute units. Inference-optimized hardware minimizes that movement, making token generation several times faster than on a typical GPU setup. In this course, the hardware you'll use is Cerebras' Wafer-Scale Engine, which is designed for fast inference by keeping the model's weights close to the compute units.
Fast inference makes lengthy agentic workflows go faster, and also unlocks latency-sensitive, real-time applications like live translation and voice agents.
Skills you'll gain:
- Compare how GPUs, TPUs, and Cerebras' Wafer-Scale Engine each handle the memory-to-compute bottleneck
- Build real-time applications powered by fast inference, including personalizing a webpage and running a multi-step workflow to analyze market signals
- Adopt concrete habits for agentic coding with fast inference, keeping your sessions focused and steering the model more effectively
My teams use Cerebras for several applications that are latency sensitive. Join and build LLM applications that respond quickly:
https://t.co/P8vchGAr22