๐ข "State-space Modeling in Long Sequence Processing: A Survey on Recurrence in the Transformer Era"
https://t.co/VaUfp3LZnl
a joint work by @TiezziMatteo , M. Casoni, @aleitteb , M.Gori (@KleineBottleM) and @StefanoMelacci
3โฃ RNNs evolutions, such as xLSTMs (@gklambauer), ODE-inspired RNNs (Neural Oscillators - @tk_rusch)
as well as novel paths for
4โฃ Learning in RNNs, such as modern evolutions of RTRL (@KhurramJaved_96) or bio-inspired methods (@BerenMillidge)
https://t.co/VaUfp3LZnl
We described several trends:
1๏ธโฃ #Transformers embraced Recurrence exploiting a stateful recurrent representation:
*๏ธโฃ all the latest models RWKV4/5/6 by @BlinkDL_AI
*๏ธโฃ #RetNet by @donglixp et al.
are competitive with quadratic self-attention!
https://t.co/VaUfp3LZnl
๐งInterested in the details of the latest SOTA architectures (#RWKV , #RetNet, #Mamba, #Griffin) for long sequences?
๐ข "State-space Modeling in Long Sequence Processing: A Survey on Recurrence in the Transformer Era" https://t.co/VaUfp3LZnl
๐จ1 month left until the #CoLLAs2024 paper submission deadline!๐จ
Donโt miss your opportunity to contribute to the field of lifelong learning. Finalize your research and submit by 15 Feb 2024!
๐ Prepare your submission: https://t.co/xLC1mbEJrS
๐๏ธ Abstract deadline: 09 Feb 2024
It is sad to lose the DeepMind office in Edmonton to the Tech layoffs and looming recession. But AI is not going away, and I am more focused than ever on the Alberta Plan for AI research.
https://t.co/DQBl7vmGaO
๐ฃ Last week I presented "Foveated Neural Computation" @ECMLPKDD
โก๏ธ Paper: https://t.co/3cOHoXxLzR
โก๏ธ Code: https://t.co/qxL3EEnOBC
A thread๐๐งต
In some visual tasks, such as image classification, salient information is distributed in regions of limited size ๐ง
1/n
(8/n) @TiezziMatteo, @MarulloSimone, Lapo Faggi, @enricomeloni_ai, @aleitteb and @StefanoMelacci investigated the effectiveness of gravitational models in an open-set class-incremental setting, by imposing stochastic coherence along the visual scanpath. https://t.co/A39BPXgHBT