Toward Linearly Regularizing the Geometric Bottleneck of Linear Generalized Attention

Published in Transactions on Machine Learning Research, 2025

Linear generalized attention mechanisms suffer from a geometric bottleneck that limits representational capacity and downstream robustness. This work provides a theoretical analysis of the bottleneck through the lens of information geometry and proposes linear regularization strategies to mitigate it. By introducing structured constraints on the attention feature maps, the approach improves both the expressiveness and stability of linear attention models, achieving consistent gains on language modeling and classification benchmarks while maintaining the computational efficiency that linear attention is designed to offer.