Extent to which dynamic latent tokenization mitigates encoding biases
Determine to what extent dynamic latent tokenization in Bolmo-like Latent Tokenizer Language Models mitigates the Latin-centric bias introduced by using UTF-8 bytes as atomic units, and quantify how strongly such models inherit biases from their underlying encoding across languages and scripts.
References
We believe that the dynamic latent tokenization can to some extent 'amortize' over the choice of the atomic unit, but it is not clear to what extent this is possible, and in how far LTLMs inherit the biases from their underlying encoding.
Our results do not settle this: a shared byte-level interface removes disjoint token identities, but whether knowledge then generalizes across scripts remains untested, and Ceq} offers a direct way to test it.