Rays as Pixels is accepted to ICML 2026 🚀
We represent camera rays as pixels and learn video + camera trajectories jointly. One model does pose estimation and view synthesis, self-consistently.
Presenting at the CVPR Video World Models workshop, June 3 AM. In Denver? Let's talk.
1/ 🚀 We’re excited to share Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation!
Tuna-2 is a native unified multimodal model that supports visual understanding, text-to-image generation, and image editing directly from pixel embeddings. 🐟✨
📄 Paper: https://t.co/eUm0tmCVHJ
🌐 Project: https://t.co/PqERmbMJuG
💻 Code: https://t.co/q3q2Kug7xN
Most unified multimodal models still rely on pretrained vision encoders, which add architectural complexity and can create representation mismatches between understanding and generation.
Tuna-2 asks a simple question: Do we still need vision encoders? 👀
Our answer is No! Tuna-2 has a completely encoder-free architecture, where images are processed directly by a unified transformer together with text tokens.
Take a glimpse at what our model can generate ↓ 🎨🖼️
Vecglypher is accepted to #CVPR 26! This work was done during my internship at Meta.
Paper: https://t.co/qLgTIhqL1j
One surprising limitation we found in today’s strongest proprietary LLMs:
Introducing Kaleido💮 from @AIatMeta — a universal generative neural rendering engine for photorealistic, unified object and scene view synthesis.
Kaleido is built on a simple but powerful design philosophy:
3D perception is a form of visual common sense.
Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models.
Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures.
It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding.
Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable.
Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings.
We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details.
🔗 Explore more results and paper: https://t.co/fOcssVKQiW
📢 Hiring a PhD intern in Generative AI @Meta (London) to work with my team on video generation.
If you’ve worked on diffusion / flow matching (plus ideally LLMs/VLMs) and have strong engineering skills, apply here: https://t.co/HxYgEVBa7T
24-week internship preferred
we just finalized a funding round to build the first virtual ai/ml engineer AdaL.
Looking for founding full-stack, AI/ML engineer, growth/sales hacker.
Company based in SF city.
Comment or DM if interested!
https://t.co/iE8YOWFW1X
Excited to share that Leffa has been accepted to CVPR 2025 🚀🚀🚀
Paper: https://t.co/ci5fZfiqat
Code: https://t.co/KMCN2gMGJY
Demo: https://t.co/UPzVrwUymr
Model: https://t.co/cw5DNODreJ
i'm comically impressed that people are coping on deepseek by spewing bizarre conspiracy theories -- despite deepseek open-sourcing and writing some of the most detail oriented papers ever.
read. replicate. compete.
don't be salty, just makes you look incompetent.