(1/4) The Algonauts Project 2025 challenge is now live!
Participate and build computational models that best predict how the human brain responds to multimodal movies!
Submission deadline: 13th of July.
#algonauts2025#NeuroAI
https://t.co/SZla1GmBrZ
1/3 Today, an anecdote shared by an invited speaker at #NeurIPS2024 left many Chinese scholars, myself included, feeling uncomfortable. As a community, I believe we should take a moment to reflect on why such remarks in public discourse can be offensive and harmful.
Vision for human-level intelligence by Yann LeCun.
https://t.co/anRdo0mOuY
To reach human-level intelligence in artificial systems, we need to first know how the brain works - there must be a simple evolved mechanism that we must be able to emulate.
Do Vision-Language Models represent space, and how?
Spatial terms like "left" or "right" may not be enough to match images with spatial descriptions, as we often overlook the different frames of reference (FoR) used by speakers and listeners. See Figure 1 for examples!
Introducing the COnsistent Multilingual Frame Of Reference Test (COMFORT), an evaluation protocol to assess the spatial reasoning capabilities of VLMs. COMFORT includes systematically designed datasets and metrics that evaluate model performance, and their deeper linguistic competence, specifically the spatial knowledge encoded in their internal representations. Find out more in the video teaser!
Almost all VLMs prefer the egocentric relative FoR with reflected transform, similar to English. Yet, we reveal significant shortcomings of VLMs: notably, the models (1) exhibit poor robustness and consistency, (2) lack the flexibility to accommodate multiple FoRs, and (3) fail to adhere to language-specific or culture-specific conventions in cross-lingual tests, as English tends to dominate other languages.
A shortened version will appear in Pluralistic Alignment Workshop @pluralistic_ai #NeurIPS2024.
It seems that the ArXiv moderators put it on hold and are eager to give it a thorough read first🤣!
So here is the Paper/Code/Data: https://t.co/0t9UGe8S7F
This collaboration turns out to be amazing, jointly led by @zheyuanzhang99, @Hu_FY_ @jayjunleee, with so many contributions and insights from @fredahshi, @Kordjamshidi@SLED_AI.
With a growing effort to align vision-language models with human cognitive intuitions, we call for more attention to the ambiguous nature and cross-cultural diversity of spatial reasoning!
Introducing REPA! We show that learning high-quality representations in diffusion transformers is crucial for boosting generation performance. With REPA, we speed up SiT training by 17.5x (without CFG) and achieve state-of-the-art FID = 1.42 using CFG with the guidance interval. 🧵[1/7]
https://t.co/MoN9vjZ93U
Excited to share a new project! 🎉🎉
https://t.co/A36dZL55ar
How do we navigate between brain states when we switch tasks? Are dynamics driven by control, or passive decay of the prev task?
To answer, we compare high-dim linear dynamical systems fit to EEG and RNNs🌀
⏬
@AnthesDaniel Third, in a collaboration w Cichy Lab, @Singer_Johannes and @martisamuser used ANNs to model different mechanisms of human task-dependent readout from ventral stream ROIs.
A unique approach imo https://t.co/FRqE0L0cZJ
Friday, August 9, 2024, 11:15 am – 1:15 pm, Johnson Ice Rink
Delighted our latest finding! We discovered that abstract representations emerge in the human hippocampus when learning to perform inference. This change in neural geometry is due to disentanglement of discovered latent and observable variables. @Nature https://t.co/Kr7ClRd95I
Preprint alert 🚨I am excited about our new paper titled “The representational nature of spatio-temporal recurrent processing in visual object recognition.” 🥳🌟
https://t.co/5wJJDxUl7V
Excited to share our new #NeurIPS2023 paper, where we introduce a simple and efficient approach to investigate how class representations emerge in vision transformers trained for image classification.
📜Paper: https://t.co/cwiXDSfmA6
A thread ��
This figure summarizes the landscape of topological neural network architectures on hypergraphs, simplicial, cellular, & combinatorial complexes in a unified graphical notation. Check out our paper and full repository of TNN equations for more ✨
What approach can explain the link between the brain and behavior? Shall we focus on neural circuits or manifolds? With @chrismlangdon & @MGENK, we provide a unifying perspective on manifolds and circuits, just out in @NatRevNeurosci: https://t.co/sM4xKUEZ6o. #tweeprint 👇
Predictions in brains and large language models:
Our latest and is out at Nature Human Behavior: https://t.co/YJ8FSt5mR4
By, once again, our wonderful team @c_caucheteux and @agramfort
Our paper got accepted at #CVPR2023!
(w/ @yu_takagi)
We modeled the relationship between human brain activity (early/semantic areas) and Stable Diffusion's latent representations and decoded perceptual contents from brain activity ("brain2image").
https://t.co/Sbwf4PhPWD
Exciting to share that our work won the Outstanding Paper Award at #NeurIPS2022!! This year, awarded to only 13 out of ~10k submissions https://t.co/7VBuKZpkz7. Super thankful for all my cross-discipline collaborators from @PrincetonNeuro, @PrincetonCS, and @DeepMind 🎉🎉🎉
Can a large language model be used as a "cognitive model" - meaning, a scientific artifact that helps us reason about the emergence of complex behavior and abstract representations in the human mind? My answer is YES.
Why and under what conditions? 🧵