What we're looking for:
• A track record of publishing at top-tier ML conferences and journals
• Hands-on experience across every stage of a research project
If this sounds interesting:
Know someone who'd be a great fit? Tag them or share this post. 🙏
Our MLR team is looking for someone who's excited about understanding, challenging, and re-imagining machine learning, pushing the boundaries of what we know and turning that knowledge into new capabilities across Apple's products.
Papers capture results. What they don't capture is the immense effort this amazing team has invested, the countless brainstorms, late-night debugging sessions, and honest debates that shaped every decision. I am grateful for being part of this team, you are amazing.
#NeurIPS2025 Mixing different datasets to train your LLM?
✨ We can help you find the perfect blend!
📈 Few small-model experiments → scaling law fit → your optimal mixture.
🎯 Easy + efficient.
Chat with us 💬 Poster #3414. Thu, Dec 4, 11am
Our new paper “𝗣𝗮𝗿𝗮𝗥𝗡𝗡: 𝗨𝗻𝗹𝗼𝗰𝗸𝗶𝗻𝗴 𝗣𝗮𝗿𝗮𝗹𝗹𝗲𝗹 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝗼𝗳 𝗡𝗼𝗻𝗹𝗶𝗻𝗲𝗮𝗿 𝗥𝗡𝗡𝘀 𝗳𝗼𝗿 𝗟𝗟𝗠𝘀” is out
📄 Paper: https://t.co/QSj2pzs0ML
💻 Code: https://t.co/J9PH1ZgaGh
WIth @FedericoDa40495@prlz77@XaviSuauC Miguel Sarabia
𝗣𝗮𝗿𝗮𝗥𝗡𝗡: 𝗨𝗻𝗹𝗼𝗰𝗸𝗶𝗻𝗴 𝗣𝗮𝗿𝗮𝗹𝗹𝗲𝗹 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝗼𝗳 𝗡𝗼𝗻𝗹𝗶𝗻𝗲𝗮𝗿 𝗥𝗡𝗡𝘀 𝗳𝗼𝗿 𝗟𝗟𝗠𝘀
For years, we’ve given RNNs for doomed, and looked at Transformer as 𝘁𝗵𝗲 LLM—but we just needed better math
📄https://t.co/lFQrUEfEvZ
💻https://t.co/Lg7gbcwgFU
Check out our new work on conditioning pre-trained generative models via activation steering. LineAS has been accepted at NeurIPS 2025. Code and paper are online:
💻 https://t.co/MJSZHe9lq0
📄 https://t.co/5owafAn77X
🚀 Excited to share LinEAS, our new activation steering method accepted at NeurIPS 2025! It approximates optimal transport maps e2e to precisely guide 🧭 activations achieving finer control 🎚️ with ✨ less than 32 ✨ prompts!
💻https://t.co/IdZOpwtFXC
📄https://t.co/sfPHk5sT2B
Our latest work on understanding Mamba models will be presented at #icml2025! Great work by last summer intern @TeresaNHuang together with Miguel Sarabia,
@amoudgl, @prlz77, and Federico Danieli.
Is the mystery behind the performance of Mamba🐍 keeping you awake at night? We got you covered! Our ICML2025 paper demystifies input selectivity in Mamba from the lens of approximation power, long-term memory, and associative recall capacity.
https://t.co/dWDYyIWLzt
Excited to share code & models for FastVLM — our blazing-fast Vision-Language Model appearing at #CVPR2025
Run it on-device with inference code optimized for Apple Silicon using #mlx.
Code: https://t.co/zrYytwr9N1
Updated paper & results coming soon. Stay tuned! 👀
New blog post that explains our work on Controlling Diffusion and LLMs using steering and optimal transport:
https://t.co/RTZyeZkIJc
This work will be presented at ICLR2025 in Singapore. See you there!
Our work on fine-grained control of LLMs and diffusion models via Activation Transport will be presented @iclr_conf as spotlight✨Check out our new blog post https://t.co/dAJQtcETNX
🚀 We're hiring an ML Researcher! 🚀
If you're an expert in LLM alignment & personalization and want to work on a world-class research team, apply here 👉 https://t.co/dODHx2IHeF
Know someone who’d be a great fit? Tag them! #MachineLearning#AI#Apple
When does composition of diffusion models “work”? Prior work (Du et al., 2023; Liu et al., 2022) has shown that composition via linear score combination can sometimes compose concepts like “dog” and “oil painting”, but why? Does it always work? https://t.co/667SNBXKVa
🚨 One question that has always intrigued me is the role of different ways to increase a model's capacity: parameters, parallelizable compute, or sequential compute?
We explored this through the lens of MoEs:
Two days until our @NeurIPS workshop on Model Interventions on Foundation Models on Sunday!!
Our second talk features @viegasf, Professor @Harvard and co-lead of @Google's PAIR initiative presenting on user-centric interpretability.
https://t.co/eD0z0pcVxS
Interested in how interpretability can be used for alignment? Join us at the workshop on foundation model interventions the 15th of December at @NeurIPSConf!
Thrilled to share the latest work from our team at @Apple where we achieve interpretable and fine-grained control of LLMs and Diffusion models via Activation Transport 🔥
📄 https://t.co/TYlwxarrWx
🛠️ https://t.co/gciUcwRNqd
1/9 🧵
I’m thrilled to announce 3 #internship openings @Apple ML Research in beautiful ☀️ #Barcelona ☀️ for 2025! Two internships on Generative Models (GM), Controllability, Interpretability, and Model Editing; and one on GM &🔈Spatial Audio. Apply: https://t.co/RG6OobIvL3
Details 🧵
I’m thrilled to announce 3 #internship openings @Apple ML Research in beautiful ☀️ #Barcelona ☀️ for 2025! Two internships on Generative Models (GM), Controllability, Interpretability, and Model Editing; and one on GM &🔈Spatial Audio. Apply: https://t.co/RG6OobIvL3
Details 🧵
Enjoy attention? Want to make it ~18% faster? Try out Sigmoid Attention. We replace the traditional softmax in attention with a sigmoid and a constant (not learned) scalar bias based on the sequence length.
Paper: https://t.co/49FeKwLAE0
Code: https://t.co/b0cp49pXKX
This was an amazing collaboration with some great researchers at Apple: Federico Danieli, @EeshanDhekane, @FlorisWeers, @danbusbridge, @PierreAblin, Tatiana Likhomanenko, Jagrit Digani, Zijin Gu, @AmitisShidani1 and Russ Webb.
More details in 🧵below.