It's my last day as an intern on @modal's training team.
With the rise of RL, it's becoming increasingly critical to have robust, reliable observability. I wanted to share some of my work building an observability system for the Modal Training Gym (https://t.co/BGuTWqZQXv), built on top of @slime_framework & @radixark miles.
Thank you @qjoyliu, @peywalt, training team, and Modal for an awesome summer!
Today we're launching Miles v0.1, an open-source RL framework for LLMs and multimodal models.
RL training is easy to start and hard to debug. Miles helps you ensure your run is correct, use hardware efficiently, and keep RL running at scale.
Over the past 9 months, 72 contributors have landed 1,326 commits, 85 GPU E2E CI tests, battle-testing Miles on frontier open models like Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, MiniMax H3, etc.
Miles powers frontier-model development and production RL workloads at @humansand, @periodiclabs, @modal, @DecagonAI, @Eigent_AI, @nebiusai, @IBM and more, on both @NVIDIAAI and @AIatAMD hardware.
Here is what we built, and why teams picked Miles🧵
It's my last week @modal.
This summer, I had the chance to work alongside @waltergoat, @colin_weld, @paulgb on our Sandbox product. It was my first time working on systems at this scale, and I couldn't be more grateful for their mentorship.
Sandboxes are full of gnarly distributed systems problems. I wrote about one of them - distributed locking - on my blog: https://t.co/ooYMVJhHqu
To ship it, I needed to break rules for practicality and performance. This led me to formally verify the design with TLA+ to catch the many race conditions.
This is also my first blog post! I don't share my writing often, but I'd like to do it more. Hope you enjoy.
🧵I post-trained Qwen3-4B-Instruct to play MegaGem, @JaneStreetGroup's new auction game.
It ranks #1 on 3-player MegaGem, above models like GPT-5.5, Claude Opus 4.8, and Gemini 3 Flash, but statistically unresolved with Gemini 3.1 Pro.
However, self-play RL wasn’t what got it there.
Today was my last day at Modal.
I had an amazing 6 months on the training team, working on sandbox infra (RDMA) and on-policy distillation methods.
Currently feeling quite ambivalent about it: will be missing the people here but also excited to start college at CMU this fall!
Applications are now open for the 2026 Paradigm Fellowship: a 4-day retreat for young people who are obsessively good at something technical.
For our fourth year, we're expanding to welcome builders across every frontier — AI, robotics, energy, bio, prediction markets, or something we haven't thought of.
Last year's cohort came from 10 countries. Some were undergrads, some were dropouts, some were founders, some came from OpenAI, SpaceX, Citadel, and Kalshi.
The format is simple: firesides, whiteboarding sessions, and time to hack. What makes it special is what happens in between, and after. Fellows have met cofounders, started companies, and gone on to raise from Paradigm and others.
I was a fellow before joining Paradigm, the retreat was a transformative trip for me, and I met some of my closest friends through the program.
Apply by June 8th. Retreat runs August 12–15th.