Wow. Recreating the Shawshank Redemption prison in 3D from a single video, in real time (!)
Just read the MASt3R-SLAM paper and it's pretty neat. These folks basically built a real-time dense SLAM system on top of MASt3R, which is a transformer-based neural network that can do 3d reconstruction and localization from uncalibrated image pairs.
The cool part is they don't need a fixed camera model -- it just works with arbitrary cameras -- think different focal lengths, sensor sizes, even handling zooming in video (FMV drone video anyone?!). If you've done photogrammetry or played with NeRFs you know that is a HUGE deal.
They've solved some tricky problems like efficient point matching and tracking, plus they've figured out how to fuse point clouds and handle loop closures in real-time.
Their system runs at about 15 FPS on a 4090 and produces both camera poses and dense geometry. When they know the camera calibration, they get SOTA results across several benchmarks, but even without calibration, they still perform well.
What's interesting is the approach -- most recent SLAM work has built on DROID-SLAM's architecture, but these folks went a different direction by leveraging a strong 3D reconstruction prior. Seems to give them more coherent geometry, which makes sense since that's what MASt3R was designed for.
For anyone who cares about monocular SLAM and 3D reconstruction, this feels like a significant step toward plug-and-play dense SLAM without calibration headaches -- perfect for drones, robots, AR/VR -- the works!
This paper is even more insane to read than the thread. Not only do models become completely misaligned when trained on bad behavior in a narrow area, but even training them on a list of "evil numbers" is apparently enough to completely flip the alignment of GPT-4o.
What if an AI agent makes a phone call, then realizes the other person is also an AI agent?
At the ElevenLabs London Hackathon, Boris Starkov and Anton Pidkuiko introduced a custom protocol that AI agents can switch into for error-proof communication that's 80% more efficient
It's mind-blowing
Wait…WHAT? Introducing Pikadditions, the easiest way to make your content stand out.
Add anyone or anything to any video, whether that’s a video you shoot yourself, or a favorite clip. Special surprise: get fifteen free Pikadditions generations when you sign up!
Go try it at pika dot art
The new Gemini Realtime API is absolutely amazing.
My new hobby is using it to annihilate at GeoGuessr...
It got within 2 blocks of the exact location on the first try 🤯
By using multiple telescopes & cameras, I can create composite images that are impossible with traditional photography, like this composition of the last quarter moon amidst the clouds.
This is formatted as a mobile wallpaper for you to use. Enjoy! I’m still posting these daily!
Newsletter: Goldman Sachs has called BS on Generative AI, and I believe that it's time that everybody follows suit - generative AI is unreliable, unsustainable, requires an entire rebuild of America's power grid, and is most decidedly not the future.
https://t.co/YULEkHYBFP
Can #LLMs truly reason over loooong context? 🤔
NoCha asks LLMs to verify claims about *NEW* fictional books 🪄 📚
⛔ LLMs that solve needle-in-the-haystack (~100%) struggle on NoCha!
⛔ None of 11 tested LLMs reach human performance → 97%. The best, #GPT-4o, gets only 55.8%.
Our 6ft humanoid HumanPlus at Stanford can autonomously put on a Nike skateboard shoe, tie shoelaces, stand up and walk.
Using two transformers & dual RGB vision, it integrates two recipes of general robotics end-to-end:
- imitating humans in real world
- large-scale RL in sim
Want to see Mona Lisa do something new? AI can make it happen! 🖼️✨#MonaLisa Transform the classics and bring art to life with AI. Ready for more creative possibilities? Follow for amazing AI-generated content! @Kling_ai#KlingAI#Kling#DreamMachine#LumaDreamMachine#Gen3