Scaling Transformers to very large graphs
Determine how to adapt the Transformer architecture used for graph representation learning to scale effectively to large graphs with potentially millions of nodes while addressing the quadratic time and memory complexity of all-to-all self-attention.
References
However, this comes at the cost of quadratic complexity, and it remains an open question how to adapt the Transformer architecture to scale effectively to large graphs with potentially millions of nodes.
— k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS The Expressive Power of GraphGPS
(2604.03815 - Schouwer et al., 4 Apr 2026) in Section 1 (Introduction)
Moving toward hardware-relevant deployments raises two immediate questions: how performance degrades under finite-shot estimation of observables, and how feature generation can be scaled to much larger graphs.
— On the Expressive Power of the Transverse-Field Ising Model for Graph Learning
(2608.17750 - Djellabi et al., 18 Aug 2026) in Section 4, Discussion, paragraph “Practical constraints and limitations”