Deprecated: $wgMWOAuthSharedUserIDs=false is deprecated, set $wgMWOAuthSharedUserIDs=true, $wgMWOAuthSharedUserSource='local' instead [Called from MediaWiki\HookContainer\HookContainer::run in /var/www/html/w/includes/HookContainer/HookContainer.php at line 135] in /var/www/html/w/includes/Debug/MWDebug.php on line 372
The emergence of clusters in self-attention dynamics - MaRDI portal

The emergence of clusters in self-attention dynamics

From MaRDI portal
Publication:6510023

arXiv2305.05465MaRDI QIDQ6510023

Borjan Geshkovski, Yury Polyanskiy, Cyril Letrouit, Philippe Rigollet


Abstract: Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent. We show that particles, representing tokens, tend to cluster toward particular limiting objects as time tends to infinity. Cluster locations are determined by the initial tokens, confirming context-awareness of representations learned by Transformers. Using techniques from dynamical systems and partial differential equations, we show that the type of limiting object that emerges depends on the spectrum of the value matrix. Additionally, in the one-dimensional case we prove that the self-attention matrix converges to a low-rank Boolean matrix. The combination of these results mathematically confirms the empirical observation made by Vaswani et al. [VSP'17] that leaders appear in a sequence of tokens when processed by Transformers.




Has companion code repository: https://github.com/borjang/2023-transformers








This page was built for publication: The emergence of clusters in self-attention dynamics

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6510023)