Formal Multi-Layer Composition Theorem for Sink Preservation

Derive a rigorous multi-layer composition theorem showing how per-layer sink margins and gap perturbation bounds control attention-sink preservation through successive transformer layers, with explicit control of value vectors, residual streams, MLP transformations, and post-attention layer normalization.

Background

The main theoretical results establish a perturbation bound and saturation-transfer lemma for a single attention layer and head. Extending these results across layers requires controlling additional components of the transformer computation, including value vectors, residual connections, MLPs, and post-attention layer normalization. The paper does not provide these controls or formalize the required assumptions, and therefore leaves a rigorous multi-layer result unresolved.

References

A tight multi-layer theorem (in the spirit of~\citep{olsson2022context}'s circuit composition) is future work.

Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis  (2609.00746 - Choi et al., 1 Sep 2026) in Appendix A.4, “Multi-Layer Composition Remark”