Your cold emails can be the golden ticket to your dream PhD lab. Here’s how:
Most applicants treat cold emails like polite introductions. That’s a mistake. A PI isn’t looking for politeness. They are scanning for signal. Can this person think, contribute, and reduce my uncertainty?
Here’s a research-oriented, high-leverage approach:
1. Start with a paper, not a profile
Don’t begin with “I’m interested in your work.” Begin with a specific claim from one of their recent papers.
Identify a gap, assumption, or unexplored extension.
2. Write a 3 to 4 sentence micro-proposal
Structure it like this:
Observation: What they did
Tension: What remains unclear or limited
Idea: Your proposed extension
Method hint: How you would approach it
This signals you are already thinking like a researcher, not an applicant.
3. Attach proof of execution, not just potential
Link 1 to 2 artifacts only:
A GitHub repo
A preprint
A tight 2-page research note
Each should directly relate to the idea you pitched.
4. Use the adjacent expertise angle
Labs do not just need clones. Position yourself as someone who brings a method or perspective they do not currently have but clearly need.
5. Ask a low-friction question
Instead of “Are you accepting students?” ask:
“Would this direction align with your current priorities, or am I missing a key constraint?”
This invites engagement, not rejection.
6. Timing is strategic, not random
Email right after:
A new paper release
A conference talk
A grant announcement
You are entering when attention is already on new ideas.
7. Subject line equals hypothesis, not a request.
Good: “Extending your X paper: idea on Y limitation”
8. Keep it under 350-400 words
Constraint forces clarity. Clarity signals intelligence.
DM me “cold-email” if you also want to create highly magnetic cold-emails that turn into interview offers.
For this week's seminar, we are excited to host @shangbinfeng from University of Washington!
Date and Time: Thursday, February 26, 11:00 AM — 12:00 PM Pacific Time.
Zoom Link: https://t.co/jmz2wb8Xyn
Title: Protocols of Model Collaboration
Abstract: Human intelligence is compositional: there is no “general-purpose” human and we are all specialized in our own ways, while collaboration protocols guide diverse individuals to come together and achieve what they cannot on their own. Moving beyond single monolithic AI models, I aim to advance compositional intelligence through model collaboration, where multiple (language) models collaborate, compose, and complement each other. In this talk I discuss three specific protocols of model collaboration: Model Swarms, multiple LMs collaboratively search in the model weight space for adaptation; Sparta Alignment, multiple LMs collectively evolve via competition and combat; and Switch Generation, pretrained and aligned stages of LMs take turns to generate segments of responses to patch the tradeoffs of RL/alignment. These protocols span diverse levels of model access and information exchange, spearheading a new paradigm of AI systems featuring compositional intelligence and collaborative development.
Excited to see everyone at the seminar!
Ethan Neumann, a 27-year-old PhD candidate in my lab, devoted himself to studying #fibrolamellar carcinoma, the cancer that claimed his life last week.
Please support his father’s effort to establish an endowed research fellowship in Ethan’s name ⬇️
https://t.co/g4orF0msmb
Highlights from my #NeurIPS2025 panel talk on AI Agents for the Gemini Enterprise launch! 🚀
• 𝗗𝗲𝗲𝗽 𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵: Automated research lifecycle
• 𝗔𝗴𝗲𝗻𝘁 𝗠𝗲𝗺𝗼𝗿𝘆: Solving the forgetful agent problem
• 𝗠𝘂𝗹𝘁𝗶-𝗔𝗴𝗲𝗻𝘁 𝗦𝘆𝘀𝘁𝗲𝗺𝘀: No-code agent creation
Model collaboration talk tour continues~
Compositional intelligence. Collaborative development. Decentralized AI. By the Many.
The methods. The vision. The hot takes. The comedy.
If you are around one of these places, let's chat!
Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other approaches for a fraction of the cost.
https://t.co/JhpyWQOpBe
Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other approaches for a fraction of the cost.
https://t.co/JhpyWQOpBe
🔥Introducing #AgentFlow, a new trainable agentic system where a team of agents learns to plan and use tools in the flow of a task.
🌐https://t.co/a0KygGDYhN
📄https://t.co/vfNz5Y7pRa
AgentFlow unlocks full potential of LLMs w/ tool-use.
(And yes, our 3/7B model beats GPT-4o)👇
🧩A team of four specialized agents coordinates via shared memory:
Planner: plan reasoning & tool calls 🧭
Executor: invoke tools & actions 🛠
Verifier: check memory status ✅
Generator: produce final results ✍️
💡The Magic:
🌀💫 AgentFlow directly optimizes its Planner agent live, inside the system, using our new method, Flow-GRPO (Flow-based Group Refined Policy Optimization). This is "in-the-flow" reinforcement learning.
📊The Results:
AgentFlow (7B backbone) outperforms top baselines on 10 benchmarks, with average gains of:
+14.9% on search 🔍
+14.0% on agentic 🤖
+14.5% on math ➗
+4.1% on science 🔬
🏆It even surpasses larger-scale models like Llama-3.1-405B and GPT-4o (~200B).
Try it yourself!
🛠️Code: https://t.co/XFvTyJt3WZ
🚀Demo: https://t.co/6RDKYW2368
🤖Model: https://t.co/3IfV4rB9Be
📊Visual: https://t.co/zef0pHIiWj
💬Join our Slack: https://t.co/30Pr64Cp9a
#agentic #llms #RL #tooluse
Introducing the Model Collaboration Tour 🤖🤝
Compositional intelligence. Collaborative development. Decentralized AI. By the Many.
The methods. The vision. The hot takes. The comedy.
LA folks, join us this week!
Get your tickets by asking me to give a talk @ your lab/school!
You asked, I deliver (2 weeks late).
A blogpost about what I did over the last 18 months to sign as a research engineer at a foundation model lab.
Link and summary below
1/n
Even though there are no fundamental reasons that conclusions from 1b models couldn’t generalize to 32b models (or larger), I rarely see this happens in practice. If anyone has ideas about this, I’m all ears. I’m so eager to be wrong in this case 🥰