Run behavior of the hybrid on real data

Determine whether the increased space caused by the encoded text and the auxiliary array changes when the hybrid index is applied to real data, including parses and minimizer digests.

Background

The hybrid index replaces each character with its frequency rank among the characters that follow the preceding character, then performs almost all backward-search steps on the resulting encoded text. The encoding typically has a small, skewed alphabet, but the auxiliary array storing the unencoded preceding characters can have many more runs than the original BWT.

Synthetic experiments show that the hybrid can be faster than competing indexes at intermediate alphabet sizes and with some noise, but its compact representation is substantially larger than compressed run-length FM-indexes and CSAs. The paper identifies the behavior of the encoding and auxiliary array on real repetitive data—particularly parses and minimizer digests—as the key unresolved issue determining whether the method is practically useful.

References

Whether that changes on real data, such as parses and minimizer digests, is the main open question.

— Mixing FM-indexes and CSAs: backward search over an order-1 rank encoding  (2609.40299 - Gagie, 30 Sep 2026) in Abstract; Section 1, paragraph “What we found”