Can a video model learn correspondence from raw video, without track labels?
Our CVPR Highlight introduces Video-GMAE, which represents a video as 3D Gaussian splats moving over time, and leads to zero-shot point tracking. Visit our poster 3:30-5:30 on Sunday!
More in thread 🧵
(6/7)
There are still clear limitations: static-camera pretraining, a 256-Gaussian budget, and difficulty with fine details under large motion. Fixing these might lead to video-SSL objectives with other emergent capabilities.
@jrichardgoodman if the reporting is accurate that the FO (Joe) wants Golden there then Joe has proven that he'll throw money at the problem regardless of if they want to be there or not (re: the Jimmy situation)
@ToolmanTA@jrichardgoodman I suppose Jimmy didn't wanna come here either until Joe threw 60 mil at him. The question is whether the Giannis situation gets nearly bad, seems like he has some places in mind to go to...
@jrichardgoodman I gotchu Coach I remember these well, I'm just asking why the change from this now? I think the same idea applies now with Gui (and others).
@jrichardgoodman Were they winning games when Lamb and Jerome were playing over JK two years ago? They were in basically the same spot as they are now record wise. So why the switch up from you now coach?