"SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE"
TL;DR: injects spherical positional embeddings into pretrained diffusion transformers to generate seamless 360° panoramas and videos without any training or optimization.
Hirschorn et al., "SpheRoPE: Zero-Shot Optimization-Free 360◦ Panorama Generation with Spherical RoPE"
Periodic RoPE + additional guidance by CFG using a prompt that encourages 360 panoramas. Allows training-free & optimization-free repurposing to generate 360 panoramas.
Excited to share our new paper: SpheRoPE! 🌐
We introduce a zero-shot training-free framework that lets DiTs generate flawless 360° panorama images and videos. Prompt a world and instantly step inside it.
📄Paper: https://t.co/HrkEzm1vkY
🌎Project Page: https://t.co/5hfpEcWYCK
Excited to share our new paper: SpheRoPE! 🌐
We introduce a zero-shot training-free framework that lets DiTs generate flawless 360° panorama images and videos. Prompt a world and instantly step inside it.
📄Paper: https://t.co/HrkEzm1vkY
🌎Project Page: https://t.co/5hfpEcWYCK
SpheRoPE unlocks instant, high-fidelity explorable environments for Flux & LTX-Video.
Huge thanks to Aaron Olender, Eli Alshan, Ianir Ideses, @FritzLior and @BenaimSagie !
💡The Solution:
Instead of slow optimization loops or fine-tuning on scarce data, SpheRoPE works entirely on the fly.
It requires zero training and zero optimization, correcting the spherical geometry completely at inference time.
SpheRoPE unlocks instant, high-fidelity explorable environments for Flux & LTX-Video.
Huge thanks to Aaron Olender, Eli Alshan, Ianir Ideses, @FritzLior and @BenaimSagie !
💡The Solution:
Instead of slow optimization loops or fine-tuning on scarce 360 data, SpheRoPE works entirely on the fly.
It requires zero training and zero optimization, correcting the spherical geometry completely at inference time by changing the RoPE of the model.
Excited to share our new paper:
Multi-View Foundation Models! 🌐
We introduce a general framework to upgrade 2D foundation models (DINO, SAM, CLIP) into multi-view consistent foundational models.
📄 Paper: https://t.co/c1gWvs27fG
⚡️ Project Page + Code: https://t.co/fgYzMD3L61
Our paper:
"LaMI: Augmenting Large Language Models via Late Multi-Image Fusion"
has been selected for an Oral Presentation at #ACL2026!
LaMI boosts LLM visual commonsense by generating complementary images from a text prompt and late-fusing their evidence into the prediction
🧵
����Splatent has been accepted to #CVPR2026
Get sharp, high-quality reconstructions from diffusion latent space.
Research by Amazon Prime Video:
pape:r https://t.co/1uVSMT7qPG
project page: https://t.co/DDZMKYyn0z
Huge thanks to the amazing team @Or_Hirsch @FritzLior @omeriko_
Splatent is accepted to #CVPR2026!! 🎉🚀
Huge thanks to the incredible collaborators who made this happen: @FritzLior, @inbarhub, and the rest of the team.
Happy to share Splatent, a new research done during my internship at Amazon @PrimeVideo! 🎬
We tackle a key issue in 3D generation: getting sharp reconstructions directly from diffusion latent space.
📄 Paper: https://t.co/znrZocsEEE
🌐 Page: https://t.co/5yL8DxauJx
Excited to share CLIMP - the first fully Mamba-based contrastive vision-language model. Unlike CLIP's ViTs, Mamba's state-space formulation favors locality & smoothness—better retrieval and OOD robustness.
with @ItamarZimerman, @Eli_Schwartz and @RGiryes
https://t.co/ytWLXemteK
Segre and Hirschorn et al., "Multi-View Foundation Models"
Train an adapter to 2D foundational models like DINO/SAM/CLIP that allows turning them into "multi-view" versions. More reliable results when doing multi-view tasks.
Excited to share our new paper:
Multi-View Foundation Models! 🌐
We introduce a general framework to upgrade 2D foundation models (DINO, SAM, CLIP) into multi-view consistent foundational models.
📄 Paper: https://t.co/c1gWvs27fG
⚡️ Project Page + Code: https://t.co/fgYzMD3L61
What can you do with it?
1️⃣ Robust Matching: Features lock onto 3D points across drastic view changes.
2️⃣ Segmentation: Click once, and MV-SAM segments across all views.
3️⃣ Geometry: Estimate globally consistent surface normals directly from features.