Local minima are extremely rare in high dimensional spaces, so if you ever feel stuck in a rut it’s probably just because you aren’t considering a wide enough set of orthogonal options
@eding421 you have no idea how annoying it was I put hours into my eccv and colm reviews and half of them came back with obvious chatgpt rebuttals… like atp how do I even know ur numbers aren’t hallucinated
If reviewers feel the rebuttal is entirely AI generated, they are even less likely to raise the score.
Also, based on past author responses I think are LLM generated, they seem to argue over every minor suggestion, even when it is simple to fix and could strengthen the paper.
opensource our rebuttal skills, distilled from a lot of successful rebuttals from myself & labmates; also my experience serving as ACs for ARR.
Hope this help for your EMNLP and incoming NeurIPS reviews :)
https://t.co/LKcb9vKuPh
GPT Image 2 has been deeply unsettling to me in the best way.
Some of its outputs make it hard for me to keep using the old criterion of vision, especially the old definition of visual representation learning.
Thus, I wrote this essay as a reflection on that shift: why knowledge may be the right name for what vision once called representation, and what can be the ultimate formulation for representation learning.
(An unexpected side path: it also led me to think about the relation between knowledge and representation through the old calligraphic relation between spirit and form 😃
https://t.co/AW8Nd8s9hw
“there has been limited evidence that generative vision models have developed strong understanding capabilities.”
“Limited” or strong evidence from 2023 and 2024 🤔 👇
https://t.co/O28Ir1KMSJ
@ducha_aiki@hpirsiav10 understands the world will be able to spot most of the errors. However, they are intended to take time and be "tricky."
3) Some pairs are intended to take time to solve. This is why we use the plausible/implausible distinction, as plausible pairs would look normal at a (3/4)
@sir4K_zen@ducha_aiki@hpirsiav10 The point is that it’s supposed to be a spatial reasoning benchmark, unlike BLINK or others which test low level vision understanding. It might take you a few seconds, but you get it right because you understand 3D.
Khangaonkar et al., "Multimodal Large Language Models Cannot Spot Spatial Inconsistencies"
A benchmark for spotting spatial inconsistencies. MLLMs are still quite far from being accurate. Reminds me of various automated benchmarks that have recently been shown to be misleading.