- The paper introduces Sayre, a template-free framework that combines exemplar-guided multi-agent synthesis with error-driven generation to create 1 million document–schema–annotation training instances.
- The paper shows that Sayre-4B reaches 72.92 overall field-level F1 on UniKIE and Sayre-2B reaches 70.47, outperforming larger on-device baselines and narrowing the gap with server-scale models.
- The paper reduces field-level errors by 17.8%, especially for dense line-item tables and contract fields, while revealing that realistic handwriting synthesis remains a key challenge.
Motivation and problem setting
Key Information Extraction (KIE) converts visually rich documents into structured data, and Large Multimodal Models (LMMs) now offer end-to-end extraction that jointly models text, appearance, and layout. However, the strongest results typically come from large on-server models whose inference cost, latency, and data-transfer requirements are prohibitive in many enterprise deployments. Compact, locally deployable LMMs are an attractive alternative, but they lack sufficient KIE supervision: annotating real enterprise documents requires field values, extraction schemas, and field-level correspondences simultaneously, making large-scale collection expensive.
Existing document synthesis approaches mitigate this scarcity but depend on hand-crafted templates or simple content replacement, which scale poorly across document categories and often fail to preserve category-specific content patterns and layout conventions. The paper proposes Sayre (2607.04636), a template-free synthesis framework that generates document–schema–annotation triples from a handful of exemplar documents per category, and additionally converts real-world failure cases into hard training examples.
The Sayre framework
Sayre comprises two complementary generation pipelines.
General data generation is instantiated as a multi-agent system. Given n exemplar documents from category c, a topic agent samples a new topic conditioned on a category identifier and a persona card (1M elite personas sampled from the persona hub of Ge et al.), increasing content diversity. A VLM-based content perception agent summarizes commonly appearing fields, their semantics, and dependencies; a layout perception agent captures page size, hierarchy, and spatial distribution. A content generation agent produces structured outputs y conditioned on the content description and sampled topic; the query schema S is derived by stripping value fields from y. A document generation agent then emits HTML code integrating y with the layout description, which is rendered into the final image. For domain-specific scenarios, a schema-guided extraction agent maps y onto a predefined target schema S′, aligning synthesized supervision with downstream applications without manual templates.
Error-driven data generation targets failure modes that general synthesis misses. Failure cases are collected from two sources: testing of a model trained on general synthetic data, and an in-house production system aggregating errors across OCR-based and end-to-end pipelines; all cases are manually annotated. Following olmOCR-style practice, an agent converts each failure case into an HTML template; parsing rules extract text blocks and align them with label values via a mapping. An LLM rewrites the text into semantically similar but distinct content for de-identification, labels are updated through the mapping to maintain alignment, and the rewritten text is reinserted and re-rendered. Where fields cannot be reliably mapped, multi-model voting over several advanced LMMs supplies re-annotations. This expands a limited set of real-world failure patterns into a larger corpus of hard examples while preserving their structural difficulty.
The two pipelines together produce 1M instances, augmented in Blender with realistic optical noise simulating acquisition conditions. Implementation uses Qwen-VL-Max for perception, Qwen3-Max for content/HTML generation, and Qwen3-VL-Plus for error-driven templates; Sayre-2B and Sayre-4B fine-tune Qwen3-VL-2B/4B backbones (15K steps, DeepSpeed ZeRO-2 on 4×A800, effective batch 64, 8192-token sequences, ~1.6M-pixel max resolution).
Evaluation is on UniKIE with field-level F1 under exact-match normalization after normalization, spanning constrained-category (business transactions, public services, regulatory records) and open-category (receipts, forms, invoices, contracts) settings.
| Model |
Constrained avg. |
Open-category avg. |
Overall |
| Gemini-3-Pro (on-server) |
82.49 |
81.65 |
82.01 |
| Qwen3-VL-Plus (on-server) |
80.32 |
70.61 |
74.77 |
| Qwen3-VL-4B (foundation) |
72.97 |
57.71 |
64.25 |
| Sayre-4B |
77.36 |
69.59 |
72.92 |
| Qwen3-VL-2B (foundation) |
68.34 |
54.34 |
60.34 |
| Sayre-2B |
74.97 |
67.11 |
70.47 |
| Best other on-device (MiMo-VL-7B-RL) |
68.41 |
60.55 |
63.92 |
Three findings stand out. First, fine-tuning on Sayre data consistently improves both backbone sizes across all categories. Second, Sayre-4B ranks first among on-device LMMs in every evaluated category, and Sayre-2B surpasses substantially larger baselines such as GLM-4.1V-9B and MiniCPM-V4.5-8B despite its smaller backbone. Third, gains are largest in the open-category setting (e.g., +18.5 F1 on Receipt and +13.7 on Contract for Sayre-4B), where models must adapt to unseen schemas — evidence that the benefit is not memorization of fixed templates but improved generalization to new document scenes. Notably, Sayre-4B's overall score (72.92) approaches strong proprietary on-server systems like Qwen3-VL-Plus (74.77), indicating that data quality and coverage can narrow the gap between compact and server-scale models as much as model scale itself.
Scaling behavior
Varying the amount of synthesized training data reveals an overall upward trend for both backbones, though the shape differs by setting. Constrained-category performance improves rapidly early and then plateaus, consistent with fixed-schema documents being learnable from moderate data volumes. Open-category performance continues to improve at later stages, matching its greater schema and layout diversity. The average-gain curves show positive improvements over the foundation model throughout, strongest for the 2B model — smaller LMMs extract disproportionate benefit from scalable synthetic supervision. This scaling is achieved without any hand-crafted templates, distinguishing Sayre from prior parameterized-sampling approaches.
Error analysis
Comparing Qwen3-VL-4B against Sayre-4B at the field level shows a 17.8% reduction in total field-level errors, with both false positives and false negatives reduced — indicating genuine improvement in field localization and schema alignment rather than a conservative prediction shift. The largest reduction comes from line-item fields (item names, quantities, prices, subtotals) in dense tables with repeated rows, where the foundation model tends to miss entries or misassociate values with keys. Substantial reductions also appear on contract clauses and business identifiers (dispute-resolution authority, receipt/document numbers, payment fields). The authors conclude that the primary benefit is enhanced schema-aware extraction over complex layouts, not merely better text recognition.
Limitations and open questions
The paper concedes that current models degrade on handwritten content and mixed printed–handwritten documents, because Sayre cannot yet reliably synthesize realistic handwriting and its visual variations. Several further questions remain open: whether the observed open-category scaling trend continues beyond the evaluated data volumes; how well the framework transfers to categories lacking even a few clean exemplars; and how sensitive error-driven generation is to the quality and coverage of the curated failure corpus, given its reliance on manually annotated seeds and multi-model voting for unmappable fields.
Conclusion
Sayre demonstrates that scene-aware, template-free document synthesis — combining exemplar-guided multi-agent generation with error-driven expansion of real failures — provides effective, scalable KIE supervision for compact LMMs. With a 17.8% field-error reduction, state-of-the-art on-device results across all UniKIE categories, and continued gains from data scaling, the work positions synthetic data quality as a primary lever for practical, locally deployable document understanding, while leaving handwritten-document synthesis as its principal unresolved challenge.