Papers
Topics
Authors
Recent
Search
2000 character limit reached

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

Published 6 Jul 2026 in cs.CV | (2607.04636v1)

Abstract: Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-server Large Multimodal Models (LMMs), while compact locally deployable models lack sufficient KIE supervision. We present SAYRE, a scene-aware document synthesis framework for generating scalable KIE training data without hand-crafted template design. Given a few exemplar documents, SAYRE captures category-specific content patterns and layout conventions to synthesize document-schema-annotation triples. It further introduces error-driven generation, which expands real-world failure cases into hard training examples while preserving their structural patterns. Experiments on constrained- and open-category KIE show that SAYRE consistently improves Qwen3-VL backbones and achieves the strongest overall performance among on-device LMMs. Data scaling experiments show an overall upward trend as more synthesized data is introduced, especially for smaller models and open-category extraction. Error analysis further shows that synthesized training reduces field-level errors by improving schema-aware extraction over dense tables, business identifiers, and contract clauses. These results establish scene-aware synthesis as an effective data-centric approach for improving practical multimodal KIE.

Summary

  • The paper introduces Sayre, a template-free framework that combines exemplar-guided multi-agent synthesis with error-driven generation to create 1 million document–schema–annotation training instances.
  • The paper shows that Sayre-4B reaches 72.92 overall field-level F1 on UniKIE and Sayre-2B reaches 70.47, outperforming larger on-device baselines and narrowing the gap with server-scale models.
  • The paper reduces field-level errors by 17.8%, especially for dense line-item tables and contract fields, while revealing that realistic handwriting synthesis remains a key challenge.

Motivation and problem setting

Key Information Extraction (KIE) converts visually rich documents into structured data, and Large Multimodal Models (LMMs) now offer end-to-end extraction that jointly models text, appearance, and layout. However, the strongest results typically come from large on-server models whose inference cost, latency, and data-transfer requirements are prohibitive in many enterprise deployments. Compact, locally deployable LMMs are an attractive alternative, but they lack sufficient KIE supervision: annotating real enterprise documents requires field values, extraction schemas, and field-level correspondences simultaneously, making large-scale collection expensive.

Existing document synthesis approaches mitigate this scarcity but depend on hand-crafted templates or simple content replacement, which scale poorly across document categories and often fail to preserve category-specific content patterns and layout conventions. The paper proposes Sayre (2607.04636), a template-free synthesis framework that generates document–schema–annotation triples from a handful of exemplar documents per category, and additionally converts real-world failure cases into hard training examples.

The Sayre framework

Sayre comprises two complementary generation pipelines.

General data generation is instantiated as a multi-agent system. Given nn exemplar documents from category cc, a topic agent samples a new topic conditioned on a category identifier and a persona card (1M elite personas sampled from the persona hub of Ge et al.), increasing content diversity. A VLM-based content perception agent summarizes commonly appearing fields, their semantics, and dependencies; a layout perception agent captures page size, hierarchy, and spatial distribution. A content generation agent produces structured outputs yy conditioned on the content description and sampled topic; the query schema S\mathcal{S} is derived by stripping value fields from yy. A document generation agent then emits HTML code integrating yy with the layout description, which is rendered into the final image. For domain-specific scenarios, a schema-guided extraction agent maps yy onto a predefined target schema S′\mathcal{S}', aligning synthesized supervision with downstream applications without manual templates.

Error-driven data generation targets failure modes that general synthesis misses. Failure cases are collected from two sources: testing of a model trained on general synthetic data, and an in-house production system aggregating errors across OCR-based and end-to-end pipelines; all cases are manually annotated. Following olmOCR-style practice, an agent converts each failure case into an HTML template; parsing rules extract text blocks and align them with label values via a mapping. An LLM rewrites the text into semantically similar but distinct content for de-identification, labels are updated through the mapping to maintain alignment, and the rewritten text is reinserted and re-rendered. Where fields cannot be reliably mapped, multi-model voting over several advanced LMMs supplies re-annotations. This expands a limited set of real-world failure patterns into a larger corpus of hard examples while preserving their structural difficulty.

The two pipelines together produce 1M instances, augmented in Blender with realistic optical noise simulating acquisition conditions. Implementation uses Qwen-VL-Max for perception, Qwen3-Max for content/HTML generation, and Qwen3-VL-Plus for error-driven templates; Sayre-2B and Sayre-4B fine-tune Qwen3-VL-2B/4B backbones (15K steps, DeepSpeed ZeRO-2 on 4×A800, effective batch 64, 8192-token sequences, ~1.6M-pixel max resolution).

Overall performance

Evaluation is on UniKIE with field-level F1 under exact-match normalization after normalization, spanning constrained-category (business transactions, public services, regulatory records) and open-category (receipts, forms, invoices, contracts) settings.

Model Constrained avg. Open-category avg. Overall
Gemini-3-Pro (on-server) 82.49 81.65 82.01
Qwen3-VL-Plus (on-server) 80.32 70.61 74.77
Qwen3-VL-4B (foundation) 72.97 57.71 64.25
Sayre-4B 77.36 69.59 72.92
Qwen3-VL-2B (foundation) 68.34 54.34 60.34
Sayre-2B 74.97 67.11 70.47
Best other on-device (MiMo-VL-7B-RL) 68.41 60.55 63.92

Three findings stand out. First, fine-tuning on Sayre data consistently improves both backbone sizes across all categories. Second, Sayre-4B ranks first among on-device LMMs in every evaluated category, and Sayre-2B surpasses substantially larger baselines such as GLM-4.1V-9B and MiniCPM-V4.5-8B despite its smaller backbone. Third, gains are largest in the open-category setting (e.g., +18.5 F1 on Receipt and +13.7 on Contract for Sayre-4B), where models must adapt to unseen schemas — evidence that the benefit is not memorization of fixed templates but improved generalization to new document scenes. Notably, Sayre-4B's overall score (72.92) approaches strong proprietary on-server systems like Qwen3-VL-Plus (74.77), indicating that data quality and coverage can narrow the gap between compact and server-scale models as much as model scale itself.

Scaling behavior

Varying the amount of synthesized training data reveals an overall upward trend for both backbones, though the shape differs by setting. Constrained-category performance improves rapidly early and then plateaus, consistent with fixed-schema documents being learnable from moderate data volumes. Open-category performance continues to improve at later stages, matching its greater schema and layout diversity. The average-gain curves show positive improvements over the foundation model throughout, strongest for the 2B model — smaller LMMs extract disproportionate benefit from scalable synthetic supervision. This scaling is achieved without any hand-crafted templates, distinguishing Sayre from prior parameterized-sampling approaches.

Error analysis

Comparing Qwen3-VL-4B against Sayre-4B at the field level shows a 17.8% reduction in total field-level errors, with both false positives and false negatives reduced — indicating genuine improvement in field localization and schema alignment rather than a conservative prediction shift. The largest reduction comes from line-item fields (item names, quantities, prices, subtotals) in dense tables with repeated rows, where the foundation model tends to miss entries or misassociate values with keys. Substantial reductions also appear on contract clauses and business identifiers (dispute-resolution authority, receipt/document numbers, payment fields). The authors conclude that the primary benefit is enhanced schema-aware extraction over complex layouts, not merely better text recognition.

Limitations and open questions

The paper concedes that current models degrade on handwritten content and mixed printed–handwritten documents, because Sayre cannot yet reliably synthesize realistic handwriting and its visual variations. Several further questions remain open: whether the observed open-category scaling trend continues beyond the evaluated data volumes; how well the framework transfers to categories lacking even a few clean exemplars; and how sensitive error-driven generation is to the quality and coverage of the curated failure corpus, given its reliance on manually annotated seeds and multi-model voting for unmappable fields.

Conclusion

Sayre demonstrates that scene-aware, template-free document synthesis — combining exemplar-guided multi-agent generation with error-driven expansion of real failures — provides effective, scalable KIE supervision for compact LMMs. With a 17.8% field-error reduction, state-of-the-art on-device results across all UniKIE categories, and continued gains from data scaling, the work positions synthetic data quality as a primary lever for practical, locally deployable document understanding, while leaving handwritten-document synthesis as its principal unresolved challenge.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.