🎓 Research ScientistID: rs-transformer-attention-lipschitz-bound
Self-Attention Lipschitz & Spectral Bound
自注意力 Lipschitz 连续与谱范数
🎯Core Definition
The Lipschitz Continuity & Spectral Norm Bounds of Self-Attention mathematically bounds output perturbation sensitivity with respect to input variations in deep Transformer blocks; standard self-attention f(X)=softmax(dXWQWKTXT)XWV is a non-linear matrix-valued mapping; by analyzing matrix spectral norms and the Jacobian of the Softmax operator (which has a Lipschitz constant bounded by 1 under L∞ and N under L2), the upper bound of the Attention layer's Lipschitz constant is formally proven: Lattn≤C∥WV∥2(1+d2∥WQ∥2∥WK∥2∥X∥2); this proof mathematically justifies why Spectral Normalization and d scaling prevent deep representation collapse (Rank Collapse) and exploding gradients.
💡Use Cases
AI research scientist deep learning theory rounds, mathematical stability proofs for ultra-deep Transformers, and adversarial perturbation bounds.
⚡Key Problems Solved
Deep multi-hundred layer Transformers suffer exponential perturbation divergence without Lipschitz bounds; this derivation establishes the theoretical invariants for LayerNorm and residual scaling.
🎯5 High-Frequency Exam Points
1
Derive the Jacobian matrix of Softmax J(z)=diag(p)−ppT and prove its operator spectral norm is bounded by 1?
2
Why does omitting d scale down Softmax Jacobian gradients to near-zero as dimension d grows, inducing vanishing gradients?
3
Explain how multi-layer pure self-attention without residual connections causes exponential Rank Collapse towards rank-1 token uniformity?
4
Contrast Pre-LN vs Post-LN through Lipschitz gradient norm backpropagation stability at initialization?
5
Explain how Spectral Normalization uses Power Iteration to enforce σmax(W)≤1 and regularize Lipschitz bounds in deep networks?
🔗Foundational Prerequisite Cards (Click to Review)