Variance and concentration of the number of distinct substrings

Estimate the variance of the number of distinct consecutive substrings in an i.i.d. alphabet string to gain deeper insight into the concentration of this count around its expected value.

Background

The paper studies the random variable D, which counts the number of distinct consecutive substrings of all lengths in an i.i.d. alphabet string, and derives lower bounds for its expectation in the binary and uniform d-letter cases. The authors explicitly identify concentration around the expectation as an unresolved issue and ask whether estimating the variance of D can provide a deeper understanding of that concentration.

References

Can we gain a deeper insight into the concentration of $D$ around $(D)$ by estimating the variance of $D$?

— The Expected Number of Distinct Substrings in an Alphabet String  (2609.19409 - Godbole, 16 Sep 2026) in Section Open Questions