Alignment of ReconSpan boundaries with linguistic structure

Determine whether the abrupt changes in ReconSpan's per-position backward reconstruction reach align systematically with linguistic structure.

Background

ReconSpan places chunk boundaries according to how far its backward decoder can reconstruct from each contextual prefix code before making an error. The appendix observes that gradual changes in reconstruction reach can result from merely advancing the endpoint, whereas abrupt drops indicate positions from which the decoder fails substantially earlier and therefore forces the greedy cover to allocate a new latent token.

The paper does not establish whether these reconstruction-driven boundary locations correspond to meaningful linguistic units or properties such as syntax, discourse structure, or information density. Establishing such an alignment could clarify what kinds of text the reconstruction criterion treats as difficult and could inform improved allocation rules.

References

Whether such drops align systematically with linguistic structure remains an open question.

ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization  (2608.12756 - Li, 13 Aug 2026) in Appendix, Section 'Boundary observations' (Appendix~\ref{app:boundary})