Attention has been the key component for most advances in LLMs, but it can’t scale to long context. Does this mean we need to find an alternative?
Presenting Titans: a new architecture with attention and a meta in-context memory that learns how to memorize at test time. Titans are more effective than Transformers and modern linear RNNs, and can effectively scale to larger than 2M context window, with better performance than ultra-large models (e.g., GPT4, Llama3-80B).
This man's legacy is mind-warping.
He invented the 99% of the future.
Then, the FBI seized his work. What are they hiding?
Here's the untold tale of a genius erased from history:
(and his most absurd inventions) 🧵