- The paper proposes a novel KG4VD method that constructs multimodal knowledge graphs from document layouts to support long-range VQA.
- It leverages an adaptive extraction loop, cross-page entity linking, and visually anchored retrieval to enhance evidence aggregation and answer synthesis.
- Experimental results show KG4VD outperforms previous methods in completeness, faithfulness, and conciseness across diverse document types.
Multimodal Graph Retrieval-Augmented Generation for Document-Level Visual Question Answering
Motivation and Limitations of Existing Approaches
Current Multimodal LLMs (MLLMs) and conventional Retrieval-Augmented Generation (RAG) pipelines are constrained by context window limitations when applied to visually rich, long documents. While recent Multimodal RAG (MMRAG) methods address these constraints via image-based retrieval, they remain fundamentally local: page- or component-level evidence retrieval is insufficient for questions demanding holistic, long-range document synthesis or compositional reasoning across disparate textual and visual evidence. Existing knowledge graph (KG) and graph-based RAG approaches are text-centric and lack robust methodology for grounding structured extractions to the visual and layout domains intrinsic to real-world documents. Most prior work on multimodal KGs is restricted to manually constructed graphs, social imagery, or unreliable automatic pipelines lacking verification, leaving automated, scalable multimodal knowledge graph construction and usage an open challenge.
KG4VD: Architecture and Methodology
KG4VD introduces a zero-shot, unified approach for the construction and utilization of Multimodal Knowledge Graphs (MMKGs) from visually rich documents. The framework addresses uneven information density, cross-modal evidence grounding, and efficient evidence aggregation and retrieval for robust question answering. The pipeline proceeds in two principal stages: offline MMKG construction and online retrieval/answer generation.

Figure 1: (a) MMRAG is effective for local VQA; (b) Text-only graph RAG supports cross-page summarization; (c) KG4VD integrates text, figures, and layout into a unified MMKG for long-range VQA.

Figure 2: Overall pipeline of KG4VD, highlighting MMKG construction, multimodal indexing, query-time retrieval, and answer generation.
Multimodal Knowledge Graph Construction
Each document page undergoes detailed structure parsing, with layout components (text, tables, figures, diagrams) detected and annotated. Instead of unreliable bounding box prediction, KG4VD employs a selection-based grounding strategy, associating each entity/relation with parser-identified regions. The adaptive extract-reflect loop, implemented via MLLM prompting, dynamically modulates extraction effort based on intrinsic page complexity. For every page, an extractor generates candidate entities/relations (add/replace/delete), a controller checks schema validity and grounding, and a reflector reviews coverage, iteratively focusing subsequent rounds on unresolved or partially covered components until graph completeness or a hard cap is reached.

Figure 3: Adaptive extraction round statistics; early rounds dominated by add operations, later rounds focus on revision and evidence recovery.

Figure 4: Distribution of extraction rounds per page, with textbooks and ESG reports requiring deeper iterations than slides or picture books.

Figure 5: Qualitative example of adaptive extraction, showing how iterative masking and revision refine the page graph.
After exhaustive page-local graph extraction, cross-page entity alignment is mediated by a cross-page judge that leverages both text and cropped image evidence to canonicalize and consolidate entities. Three categories (same, related, unrelated) are supported to robustly model coreference and inter-entity relations. The full document-level MMKG is constructed via iterative canonicalization, edge deduplication, and incorporation of visual clusters.

Figure 6: Entity and relation counts across the four DLVQA domains, demonstrating KG4VD's broader coverage versus other graph-RAG baselines.
Multimodal Indexing and Query-Time Retrieval
KG4VD employs unified GME encoders for both page-level (image) and entity-level (text + image crop) representations, supporting retrieval at both granularities. To avoid high-noise expansions from undirected graph traversal, KG4VD anchors retrieval via top-ranked page images, treating grounded entities as Personalized PageRank (PPR) seeds. The query analyzer dynamically selects the expansion mode (local, multi-hop, or document-level), thus tailoring the scope of graph exploration to information need complexity. Expansion candidates are reranked with a state-of-the-art multimodal reranker, and entity/relation/page selection budgets are mode-specific (Table: PPR configs, not shown here for brevity). Answer generation is two-stage: first a preliminary synthesis is constructed from retrieved graph elements, then final answer consolidation verifies and grounds the output in both textual and visual evidence.
DLVQA Benchmark: Comprehensive Evaluation Protocol
A key limitation of prior VQA benchmarks is lack of reference summaries and supporting fact-level annotation for document-level queries. DLVQA introduces 525 manually curated question-answer pairs over 3,441 pages from four diverse domains: technical slide decks, history textbooks, picture books, and corporate environmental reports. Each QA item includes: the question, expected answer summary guidance, a reference summary, and atomic supporting facts with explicit page and modality grounding.

Figure 7: Representative pages from DLVQA domains, illustrating the high diversity in content structure and multimodal density.
This enables fine-grained reference-based evaluation using the FineSurE metric for faithfulness, completeness, and conciseness, and enables reference-free pairwise comparisons for result diversity and empowerment analysis.
Experimental Results and Ablations
KG4VD achieves the strongest overall performance across all evaluated document-level VQA settings. On MMLongBench-Doc with 2,000 packed pages, it surpasses MegaRAG (44.52% vs. 41.18% accuracy) and leading multimodal retrievers (e.g., ColQwen, GME). On cross-domain DLVQA, KG4VD attains the highest FineSurE overall score (60.3), exceeding MegaRAG (+2.89), GME (+4.1), and graph-based text-only methods. Notably, KG4VD attains the greatest completeness (42.45), directly reflecting an advantage in long-range, multi-page evidence aggregation. Faithfulness (82.43) and conciseness (56.01) also outperform or match state-of-the-art.
Ablation analysis reveals:
- Absence of adaptive, multi-round extraction reduces completeness significantly.
- Disabling cross-page entity linking damages completeness by over 5 points.
- Direct entity retrieval without page anchoring produces the largest drop, signifying the necessity of visually guided graph traversal.
- Non-adaptive, fixed-range graph diffusion (i.e., disabling query-adaptive retrieval) primarily hampers conciseness, confirming the need for nuanced, intention-aware expansion.
Domain-wise analysis (not shown here, see Appendix of the paper) shows KG4VD's adaptability: it dominates on technical slides and visually complex reports, and is highly competitive even on text-dense textbooks.
Theoretical and Practical Implications
KG4VD unifies the strengths of page-centric image retrieval for visual context with large-scale multimodal graph reasoning, supporting both granular and holistic evidence composition. Its adaptive loop for page extraction, coverage-driven masking, and MLLM-based cross-page linking constitutes a step toward scalable, verifiable, and reusable MMKGs. The separable construction and retrieval phases enable robust reuse and online efficiency for new queries.
Pragmatically, KG4VD demonstrates that iterative, layout-grounded, and visually anchored graph-based RAG is tractable and scalable (see computational analysis in the Appendix), enabling high-quality long-range VQA for real-world, multimodal corpora in scientific, educational, reporting, and narrative domains. DLVQA sets a new standard for evaluation, supporting comprehensive, reference-backed analyses of model output faithfulness and evidence completeness.
On the theoretical side, decoupling knowledge base construction from retrieval/generation enables improved interpretability, fact verification, and foundation for future explainable AI work. The use of multimodal entity resolution via both textual and visual evidence is a critical advance over naively text-centric entity matching and supports robust cross-page coreference.
Future Directions
Further work is needed to extend KG4VD's MMKG construction to higher-order modalities (audio/video/interactive content), and to integrate explicit multi-step reasoning algorithms (e.g., Think-on-Graph extensions) for more sophisticated evidence path synthesis. Efficient graph expansion and retrieval for very large collections will also benefit from additional scaling research in indexing and online incremental updates. Improved visual entity alignment and canonical object understanding, potentially leveraging recent advances in open-vocabulary detection, will close remaining gaps, especially for highly abstracted or artistic documents.
Conclusion
KG4VD represents a comprehensive framework for multimodal document understanding, tightly coupling adaptive, grounded MMKG construction, cross-modal alignment, and dynamic, visually guided retrieval to meet the demands of true document-level VQA. Empirical evidence indicates robust advantages in faithfulness, completeness, and conciseness over both multimodal and text-centric graph RAG baselines on demanding, long-context evaluation suites. This paves the way for new research into transparent, traceable, and scalable multimodal information access.