β’ TL;DR: πππππ¬ ππ¨ π§π¨π π§πππ ππ¨ π₯π¨π¨π€ ππ―ππ«π²π°π‘ππ«π: Gaze Attention focuses only on what matters, matching or outperforming dense attention with up to 90% fewer visual KV entries.
β’ Project Page: https://t.co/HHjiCqyiD1
Computer vision researcher? π―ππ ππππ. In this video, you'll see that RL makes MLLMs see better than SFT.
Paper: https://t.co/pwJN8C1HSH
Venue: ICLR 2026
Project Page: https://t.co/6P65bImapb
RL makes MLLMs see better than SFT
New research by NAVER AI Lab & KAIST shows that Reinforcement Learning fundamentally reshapes MLLMs' vision encoders. RL leads to stronger, precisely localized visual representations, boosting performance on vision-related tasks & even outperforming larger models!
So we took it one step further:
If RL reshapes the vision encoder, can we πππππ it to build better MLLMs?
β Yes. Our results show, for example, πΉπ³-πππππππ SigLIP1 > SigLIP2 (with much lower relative training cost).
Please, check out our project page.
4/4
Microsoft is releasing Github Copilot X π
It includes:
β’ AI-generated answers from code docs
β’ Chat interface for code suggestions
β’ Copilot for the command line
β’ Voice interface with Copilot
β’ Copilot for pull requests
Okay, NOW it's so over.
Thereβs zero doubt that GPT-4 cannot solve robotics. Suppose GPT-4 is multimodal - robot control signal is just another modality right? What makes it so special? 3 reasons:
Data. Data. Data.
Image, video, audio, text are abundant online. Robot control data is not even close.
helpful links i am aware of for trending projects:
1. papers: https://t.co/24A4szwikY
2. papers+code: https://t.co/IuT0OdvrGu
3. code: https://t.co/JFOm6LgjsP