Reliable Identification of Corrupted Tokens

Develop reliable token-level mechanisms for identifying which tokens have been corrupted in autoregressive TokenCom transmissions, enabling detected errors to be concealed through contextually consistent regeneration and limiting error propagation through generated sequences.

Background

Autoregressive large models can propagate a single corrupted token through the remainder of a generated sequence, especially when the corrupted token occupies an influential position in a hierarchical or coarse-to-fine tokenization order. Although a receiver with a generative model can regenerate a detected error, regeneration is useful only when the system can distinguish corrupted tokens from valid ones. The paper explicitly characterizes this identification task as the central unresolved problem.

References

The open problem is then less the correction itself than reliably identifying which tokens are corrupted.

From Semantic to Token Communication: The Next Paradigm for Large-Model-Driven 6G Intelligent Connectivity  (2609.10714 - Ma et al., 9 Sep 2026) in Section 6.2.5, Token Error Propagation