Explain the source-domain reversal across training granularities

Determine why In-Domain Synthetic reconstruction achieves lower held-out character error than Out-of-Domain Synthetic reconstruction under page-level training but performs substantially worse under crop-level training.

Background

The paper compares two synthetic-data sources for Thai OCR: In-Domain Synthetic pages reconstructed from real Thai printed documents and Out-of-Domain Synthetic pages reconstructed from public English documents translated into Thai. These sources are evaluated with both page-level training, in which the model sees complete document pages, and crop-level training, in which it recognizes individual detected document regions.

The preferred source domain reverses between the two training regimes. In-Domain Synthetic yields the best synthetic result under page-level training, with a held-out median character error rate of 1.82%, whereas it performs worse than Out-of-Domain Synthetic under crop-level training, with median character error rates of 15.59% and 5.52%, respectively. The authors explicitly leave the cause of this reversal unresolved.

References

The cause of this reversal remains unclear and warrants further study.

— How Far Can Synthetic Data Take Thai OCR?  (2609.03595 - Pipatanakul, 3 Sep 2026) in Section 4.3, "Does the Training Granularity Determine What Transfers?" (Section \ref{sec:granularity})