Excited to present our #TransNeXt poster at #CVPR! 🖼️ We'll be showcasing it from 17:15 to 18:45 at Poster Session 4 in Exhibit Hall (Arch 4A-E), Poster #307. Stop by to read our poster, and feel free to email me if you want to discuss more! #CVPR2024#ComputerVision#AI#ML
@Birchlabs In addition, this article provides a better non-linear down-sampling method, which can make this kind of attention have linear complexity when processing variable-length feature maps. Its multi-scale extrapolation is also quite significant. I hope it brings some inspiration (2/2)
@Birchlabs The working method looks like the pixel-focused attention in this work: https://t.co/5jbtVeYRRk. Except that the processing of natten's edge pixels is different from the traditional sliding window attention. This work uses query-centered window to ensure translation equivariance.
@Birchlabs This work uses the smallest form of 3x3 window to reduce the area of repeated processing. The ablation experiment also shows that when there is a downsampled global feature map, a larger window may not necessarily improve performance. (1/2)