LVLM Abilities for Long-Context Document Understanding
Establish the capabilities of Large Vision-Language Models (LVLMs) for long-context document understanding by determining whether these models can reliably understand and answer questions over lengthy, multi-page documents.
References
However, their abilities on long-context DU remain an open problem.
— MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
(2407.01523 - Ma et al., 2024) in Abstract, page 1
This leaves the cost/accuracy trade-off across VLM scale largely open.
— An Empirical Study of VLM Pipelines for Long-Document QA
(2609.29933 - Ak et al., 24 Sep 2026) in Related Work, Section 2.2, subsection “Vision-Language Models”
Finally, our results assume English-language documents and the VLM families we tested, and whether the rankings carry over to other languages and to smaller on-device VLMs is open.
— An Empirical Study of VLM Pipelines for Long-Document QA
(2609.29933 - Ak et al., 24 Sep 2026) in Limitations