🚀 #ECCV2026 𝐈𝐧𝐭𝐫𝐨𝐝𝐮𝐜𝐢𝐧𝐠 𝟑𝟔𝟎𝐂𝐢𝐭𝐲𝐀𝐫𝐞𝐧𝐚 🌆🧭
"Can AI truly understand and navigate a real city?"
We introduce 360CityArena, a realistic urban navigation benchmark built from 360° videos of Akihabara, Tokyo. 🇯🇵
🏙️ Our arena:
✅ 602 real-world 360° videos*
✅ 85 streets*
✅ 175 navigation & spatial reasoning tasks
💡 Our mission:
Evaluate whether AI agents can understand, navigate, and reason about real-world urban environments.
The gap is still huge:
👤 Human: 77.3%
🤖 Best LMM-based agent: 17.1%
What will it take for AI to truly navigate the cities we live in? 🌏🤖
✨ The entire environment and all tasks are fully open-sourced, so anyone can build, evaluate, and explore with 360CityArena.
📄 https://t.co/GUZhpSIWL0
🌐 Project Page: https://t.co/UnhReKr9Qr
*: from 360RVW [Takenawa+] (our group's work)
#360CityArena #EmbodiedAI #MultimodalAI #UrbanNavigation
Introducing Hy3D WorldClaw——an agentic workflow that generates large scale 3D open worlds from text prompts.🚀🚀🚀
Not video, Not Gaussian Splatting, Every scene generated by WorldClaw is freely explorable and built entirely from editable, game-ready 3D assets with high-quality geometry and textures.
Project Page: https://t.co/KJXFWHLb1S
Sekai2: From World Exploration to Interactive World Modeling
TL;DR: A large-scale video dataset for interactive world models, combining long-form videos, camera trajectories, temporally grounded semantics, and revisits. Its 360° revisit-rich videos target long-term spatial memory and consistent world generation.
https://t.co/VX5dzOREqX
What if every video you ever shot could turn into a permanent 3D hologram in the exact place it happened?
I did it with photos first. Now I can do it with any video.
Here’s a mama deer and her baby deer, filmed with my iphone a few days ago -- now floating as geo anchored holograms exactly where I captured them.
We mapped an entire shopping mall - around 50,000 m² - by walking it with an Insta360 X5. 17,959 fisheye camera poses, all registered. Then we relocalized an ordinary phone video against it, frame by frame.
99.9% of frames relocalized. Median pose jitter: 1.7 cm.
STOP watching robot videos. Start this weekend:
• Read Action-to-Action Flow Matching
• Implement IK with plain gradient descent (then you will find out why nobody ships that lol)
• Simulate a scripted pick-and-place (zero learning involved) -- this is actually fun!
• Set up one MuJoCo scene from scratch, no pre-built env
• Train behavior cloning on 20 demos and measure exactly where it drifts
• Break something in sim on purpose and trace the error to the line
You have way too much to do. Bookmark & Repost.