Scalability of the Local-Replacement Approximation for Top-Down Tokenisers

Determine whether the vocabulary overlap observed between local-replacement-approximation tokenisers and exact-deletion-scoring tokenisers persists at the vocabulary sizes used in the main experiments.

Background

Top-down tokenisers use a local replacement approximation to estimate the cost of deleting a token, whereas exact deletion scoring would require recomputing the corpus objective for every candidate token at every pruning step. At small scale, the approximation produces an 81.5% vocabulary overlap with exact scoring, but the paper does not establish whether this agreement remains at the larger vocabulary sizes used in the principal experiments. The authors explicitly leave this scalability question open.

References

Whether this holds at the vocabulary sizes used in our main experiments, however, remains open.

Objective vs. Search: Decomposing What Makes a Good Tokeniser  (2609.19145 - Yavuz et al., 16 Sep 2026) in Limitations, Approximate top-down deletion scores paragraph; Appendix, Section Effect of the Top-Down Approximations, Local replacement paragraph