Effectiveness of Task-Specific Adaptation for KhatianDoc

Determine how much of the performance gap observed for multimodal large language models on the KhatianDoc benchmark can be addressed through task-specific fine-tuning or retrieval-augmented methods, compared with zero-shot evaluation.

Background

KhatianDoc evaluates six multimodal LLMs exclusively in a fixed, zero-shot setting. The benchmark does not include fine-tuned models, retrieval-augmented systems, or in-context-example conditions, so the reported failures establish a diagnostic baseline rather than an upper bound on achievable performance. The unresolved issue is whether task-specific adaptation or retrieval augmentation can substantially reduce the observed gap on Bengali handwritten land records, Ana-Ganda arithmetic, structured field extraction, and legal document question answering.

References

No fine-tuning or retrieval-augmented condition is evaluated, so we cannot say how much of the observed gap is addressable with task-specific adaptation or in-context examples.

KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records  (2609.03597 - Hasan et al., 3 Sep 2026) in Limitations, paragraph “Zero-shot evaluation only”