Determine whether sentence-length differences reflect language conventions or segmentation differences

Determine whether the greater average number of words in SaT-generated sentences than in Nupunkt-generated sentences reflects genuine language conventions or differences between the Nupunkt and SaT sentence-segmentation systems used for Institutional Books: Harvard Library Enriched Text.

Background

The pipeline uses Nupunkt for 138 languages with English-like punctuation and the SaT neural segmenter for other languages. The resulting sentence statistics show similar average character lengths but substantially different average word counts: approximately 23.3 words per Nupunkt sentence versus 36.7 words per SaT sentence.

The authors explicitly state that the cause of this discrepancy is unresolved. Clarifying whether it arises from linguistic conventions or from the behavior of the two segmentation systems would help interpret cross-language corpus statistics and evaluate the comparability of the resulting sentence units.

References

Whether this reflects language conventions or segmenter differences is unclear.

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale  (2608.19026 - Lowry-Duda et al., 19 Aug 2026) in Section 4, “Processing Pipeline,” subsection “Sentence Segmentation,” subsection “Results”