Scaling and Domain Generalization of Relation

Investigate the behavior of Relation-based decoder-only language models at substantially larger scales and in multimodal or post-training settings.

Background

The paper introduces Relation as an alternative token-mixing primitive and evaluates Full Relation, FlashRelation, Linear Relation, Hybrid Relation, and Relation Cache in decoder-only LLMs ranging from approximately 10M to 100M parameters. The evaluated experiments focus on language modeling and selected secondary benchmarks under pretraining-style settings. The authors explicitly identify the behavior of their approach beyond these experimental regimes—particularly at substantially larger model scales and in multimodal or post-training applications—as unresolved.

References

Behavior at substantially larger scales and in multimodal or post-training settings remains open.

Ask Self, Ask Others: Relation Is All You Need  (2608.20172 - Ge et al., 20 Aug 2026) in Section 7, “Limitations and Conclusion”