Effectiveness of Task-Specific Adaptation for KhatianDoc
Determine how much of the performance gap observed for multimodal large language models on the KhatianDoc benchmark can be addressed through task-specific fine-tuning or retrieval-augmented methods, compared with zero-shot evaluation.
References
No fine-tuning or retrieval-augmented condition is evaluated, so we cannot say how much of the observed gap is addressable with task-specific adaptation or in-context examples.
— KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records
(2609.03597 - Hasan et al., 3 Sep 2026) in Limitations, paragraph “Zero-shot evaluation only”