Establish cross-system stability of HGR measurement

Establish the measurement stability of the Historical Grounding Ratio across narrative types and source structures by constructing larger human-annotated datasets with information units aligned across annotators and reporting the corresponding segmentation and source-classification protocols.

Background

HGR depends on explicit information-unit segmentation and source-classification rules rather than functioning as a protocol-independent universal score. The study used machine annotation with human verification of only 11 segments, while other systems may use different categories, granularities, and evidentiary standards. The unresolved measurement problem is whether HGR remains stable and comparable across systems when annotation protocols and narrative structures differ.

References

Other systems, however, may use different source categories, granularities, and evidentiary standards. Two narratives with the same HGR may still differ substantially in the importance of their historical information, evidence quality, and narrative organization. Cross-system comparisons of HGR must therefore report both information-unit segmentation and source-classification rules. Building larger human-annotated datasets with information units aligned across annotators would help assess the measurement stability of HGR across narrative types and source structures.

Balancing Evidence and Interpretation: Historical Grounding Ratio as a Design Parameter for AI-Generated Urban Storytelling  (2608.24157 - Zhang et al., 25 Aug 2026) in Research Boundaries and Future Directions, Section 6.7