Our 2025 CVPR work of Associative Transformer is one of the studies on the global workspace in Transformer architectures, but from a computer vision perspective. We also relate to the Hopfield network to explain how long-term information is maintained in the shared space.
New Anthropic research: A global workspace in language models.
Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with.
We found a strikingly similar divide inside Claude.
This past week several people asked me about how Loop Transformers are related to Energy Transformers.
Energy Transformers are:
👉 looped transformers
👉 energy-based models (EBMs)
👉 Dense Associative Memories — generalized Hopfield nets with superior scaling laws for information storage
That combo is powerful:
👉 looping = iterative refinement = reasoning in the latent space (not in token space)
👉 energy-based = stability of token dynamics
👉 Associative Memory = strong retrieval capabilities
Put together: models that settle into good solutions, not just predict next tokens.
I’ve been especially excited about this class of models for a while. They feel like a promising direction for more stable, interpretable, and memory-rich AI systems.
This week at #ICLR2026 we are presenting NRGPT https://t.co/UkXt54Dzhj, which is a:
👉 a looped transformer
👉 a stable Dense Associative Memory
👉 works great on ListOps and real text
Original Energy Transformer paper: https://t.co/UOfrwZXtOf from NeurIPS 2023
Just posted a new preprint: The Stream of Computation: Temporal Continuity as a Missing Ingredient for Artificial Consciousness. w/ @yweisun & @manuelbaltieri Not a technical paper, but more an essay on the idea of creating AI systems that are always on.
https://t.co/5OlHK6yQsR
Our paper "MCM: Multi-layer Concept Map for Efficient Concept Learning from Masked Images," has been accepted at #ICLR2025 Workshop on Deep Generative Models! 👏
MCM is an efficient concept learning method that works even with highly masked images.
https://t.co/hDJcz9cqJX
Our recent paper "Associative Transformer" has been accepted at the main conference of #CVPR2025! 🎉
https://t.co/Ur40LoZc6I
Using sparse bottleneck attention and Hopfield associative memory (hence the name).
Thanks to all my collaborators.
#CVPR2025#Transformers#hopfield
Tokyo AI Safety Conference 2025 is happening in April. Let's build a safer AI future together. Some talks from the previous event can be found here:
https://t.co/2HIM21lP4Q
https://t.co/WqeGopFIjy
Here's my conversation with @CharanRanganath, psychologist & neuroscientist at UC Davis, specializing in human memory. We talk about memory, imagination, deja vu, false memories, why we remember, why we forget, and a lot of other fascinating mysteries about the human mind.
It's here on X in full, and is up on YouTube, Spotify, and everywhere else. Links in comment.
Timestamps:
0:00 - Introduction
1:03 - Experiencing self vs remembering self
14:44 - Creating memories
24:16 - Why we forget
31:53 - Training memory
42:22 - Memory hacks
54:10 - Imagination vs memory
1:03:29 - Memory competitions
1:13:18 - Science of memory
1:28:33 - Discoveries
1:39:37 - Deja vu
1:44:54 - False memories
2:04:59 - False confessions
2:08:45 - Heartbreak
2:16:19 - Nature of time
2:24:00 - Brain–computer interface (BCI)
2:38:04 - AI and memory
2:48:18 - ADHD
2:55:15 - Music
3:05:00 - Human mind
New blog post: Continual Learning, Compositionality, and Modularity - Part III
Human reasoning identifies abstract patterns from few examples and generalizes them to new inputs through compositional and continual knowledge learning.
https://t.co/HE0EfWdGQ0
#AI#LifelongLearning
To alleviate forgetting in continual learning, our team recently proposed Remembering Transformer🧠, inspired by the brain's Complementary Learning Systems. Initial results demonstrate new SOTA performances in class-incremental learning tasks.
Remembering Transformer utilizes a mixture-of-adapters architecture and a generative model-based novelty detection mechanism to dynamically route samples to relevant adapters.
Uncertain neurosymbolic AI is the definitive solution. LLMs possess a substantial amount of world knowledge, yet they lack a framework for stable modularization of knowledge in an online learning context, if the brain indeed operates on the premise of Global Workspace.
#AAAI24
Registration is open!
If you are fascinated by the mysteries of consciousness, join us at ASSC27 in Japan. This is a perfect opportunity to meet with consciousness researchers and visit Tokyo.Abstract submission deadline is Feb 23.
https://t.co/ROdoGJzXnN