Does 3D reconstruction have to be complex?
We answer this question with PointDiT (#ICML2026): a minimalist pixel-space Diffusion Transformer without bells and whistles.
We show that a plain ViT can estimate dense 3D point maps by operating directly on raw patches. No hybrid ViT+Conv architectures, no lossy VAEs, no complicated training losses. (1/5) 🧵👇
MediaPipe Hand Landmarker를 테스트해 봤는데,
구글은 진짜 미친놈이 맞습니다.
CPU만으로 이 정도 성능이 나온다고?
엄청 가볍고 손 추적도 상당히 잘 됩니다.
물론 손이 겹치거나 가려지면 추적이 흔들리긴 하지만
이건 다른 모델들도 비슷한 문제라 크게 단점은 아닌 것 같네요.
정말 인상적입니다.
Robots need to feel the world to operate in it.
Most manipulation policies today are tactile-blind. They either cannot interpret high-frequency tactile signals or treat them as a static channel. And the field lacks enough touch-rich datasets to train tactile-reactive policies at scale.
T-Rex was built to answer both. @Dantong_Niu and team, advised by @DrJimFan, @drfeifei, @JitendraMalikCV, @pabbeel@trevordarrell have tested whether a robot policy can react to high-frequency tactile signals the way human hands do, without giving up the generalization power of modern VLAs.
The result: 65% average success rate across 12 real-world tasks. +30 absolute points over the strongest baseline. In one year, tactile VLAs have gone from promising to outperforming non-tactile baselines like pi0.5 on dexterous tasks.
#Robotics #SharpaWave #Sharpa #EmbodiedAI #DexterousManipulation #TactileSensing #RobotLearning
Project link: https://t.co/1D7fCFHEZV
Robots need memory to handle complex, multi-step tasks. Can we design an effective method for this?
We propose MemER, a hierarchical VLA policy that learns what visual frames to remember across multiple long-horizon tasks, enabling memory-aware manipulation.
(1/5)