New work with @nvidia: evaluating robot policies entirely inside a world model. The policy acts, the model imagines the consequences, and the imagined evals predict real-world results. 🧵
real vs world-model rollout side by side📷
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini 🚀
Today, we’re sharing the @GoogleDeepMind white paper for GE 2, our first native multimodal embedding model. Whether it’s text, audio, video, or image, GE 2 provides a unified representation of the input.
I trained an autoencoder that reconstructs images with zero reconstruction loss.
No MSE. No image space supervision.
The only signal: "According to you, does your output look like your input through your own eyes?"
It works.
Blog link, demo and summary 👇
@dwarkesh_sp@ericjang11 Transformer have the global context with attention, they can choose to attend or not attend to local or global information. It’s not an architecture issue but rather optimization.
Demo night, @spc NYC, May 20
Lineup
@LumaLabsAI — Ray3 (reasoning video model)
@ElevenLabs — ElevenMusic (studio-grade AI music)
@SentienceCom — personal AI “digital twins”
@sandbar — voice ring for notes/actions
@floraai — unified creative AI
Apply below
Probably the lowest piece of argument clinic in this week of @TheEconomist.
The piece poorly tries to wash away the history of Mughal rule over India in guard of giving us paneer and biryani? What? 🤷
Today, we released Lyra 2.0, a framework for generating persistent, explorable 3D worlds at scale, from NVIDIA Research.
Generating large-scale, complex environments is difficult for AI models. Current models often “forget” what spaces look like and lose track of movement over time, causing objects to shift, blur, or appear inconsistent. This prevents them from creating the reliable 3D environments required for downstream simulations. Lyra 2.0 solves these issues by:
✅ Maintaining per-frame 3D geometry to retrieve past frames and establish spatial correspondences
✅ Using self-augmented training to correct its own temporal drifting.
Lyra 2.0 turns an image into a 3D world you can walk through, look back, and drop a robot into for real-time rendering, simulation, and immersive applications.
➡️ Learn more: https://t.co/ROR7miJeCU
📄 Read the paper: https://t.co/1osU9EGjGD
Protecting interest of one group doesn’t automatically mean harming other.
He also didn’t say I will not let robotaxis operate period. Although, it will be a terrible decision.
The question is asked at 15:00
https://t.co/tOLt4xtR5Z
Congrats to NYC Mayor on yet another push to harm the majority in the interest of protecting a minority interest group
Waymo rollout in NYC has been stopped
Today, NVIDIA is launching the next paradigm shift in GPU programming: cuTile BASIC
Write perf portable BASIC kernels and deploy them at any scale from edge inference devices like your calculator to entire GPU clusters
We're going back to BASIC
https://t.co/meF2T0jUSc
My feed is talking about custom CUDA kernels, so I started digging in last week. I’ve put together a "fill in the missing code" guide for reproducing GPT-2 here: https://t.co/JI2MKPUSHl
The boilerplate & paths are wired up-you just implement the forward paths with kernels.
My first pass is on the main branch (```git checkout main```). There’s still plenty of memory and synchronization overhead to iron out for anyone looking to get closer to the hardware.
NYC is becoming a hotbed for frontier tech.
Hear from a few of the names leading the way including @josephfkrause (@RadicalAI) and Rob Cochran (@faunarobotics) on March 31st @ SPC NYC moderated by our friend @EricNewcomer.
Very interesting!
I’m confused by the fact that my timeline is filled with CVPR papers declaring spatial problems like 3d reconstruction, monocular depth estimation, etc completely solved.
Can anyone help me understand why do we still need depth cameras and prebuilt object libraries with meshes to build digital twin? What are the remaining gaps? 🧐
@lesDecroissant@sarahookr In theory the manipulation only gets easier instead of paying people you can just buy enough tokens and manipulate the result.