Papers
Topics
Authors
Recent
Search
2000 character limit reached

READoc: Unified Document Extraction Benchmark

Updated 15 July 2026
  • READoc is a unified benchmark for realistic document structured extraction that converts complex PDFs into Markdown, recovering text, hierarchy, tables, and formulas.
  • The benchmark evaluates complete documents rather than isolated tasks by integrating OCR, layout analysis, and semantic reconstruction through standardized modules.
  • Its dual-source dataset from arXiv and GitHub highlights diverse challenges, emphasizing the trade-off between global structure and local precision in document parsing.

Searching arXiv for READoc benchmark and closely related document parsing benchmarks. READoc is a unified benchmark for realistic Document Structured Extraction (DSE) that defines the task as end-to-end conversion from complete, real-world, multi-page PDF documents into semantically rich Markdown (Li et al., 2024). Its central premise is that document understanding should not be evaluated as a set of isolated subtasks such as OCR, layout analysis, table recognition, formula conversion, or reading order detection, but as a document-level structural reconstruction problem in which textual content, hierarchy, tables, formulas, and ordering must be recovered jointly (Li et al., 2024). In this formulation, PDFs are treated as the natural unstructured input and Markdown as a practical target representation that is lightweight enough for downstream processing while remaining expressive enough to encode headings, lists, tables, and formulas; tables and formulas are additionally supported with LaTeX syntax inside the Markdown variant (Li et al., 2024).

1. Conceptual reframing of document structured extraction

READoc was proposed in response to what it characterizes as fragmented and localized benchmark paradigms in DSE research (Li et al., 2024). Earlier benchmarks often measure only one capability, such as layout analysis, OCR, table recognition, formula conversion, ToC extraction, or reading order detection, and many operate at the single-page or local-region level rather than the full-document level (Li et al., 2024). This fragmentation makes cross-system comparison difficult because different systems are trained and evaluated against incompatible outputs, metrics, and task scopes (Li et al., 2024).

The benchmark therefore reframes DSE as a realistic PDF-to-Markdown conversion task. In this framing, the relevant question is not whether a system can extract text from an easy page, but whether it can reconstruct the semantic organization of a complete document, including document-wide logical structure and globally consistent reading order (Li et al., 2024). This is closely aligned with later difficulty-aware and downstream-oriented evaluation efforts. Dr. DocBench similarly argues that existing OCR and document parsing benchmarks under-measure real document intelligence, especially on expert-domain structures, long-document continuity, and reading order (Yang et al., 31 May 2026). RealDocBench likewise argues that page-level or string-level similarity scores correlate poorly with downstream needs in regulated document workflows, and instead emphasizes field-level correctness and layout understanding on real-world pages (Joshi et al., 5 Jun 2026). Together, these later developments suggest that READoc’s original unification agenda anticipated a broader shift from localized parsing metrics toward realistic document reading evaluation.

A further implication is that READoc occupies an intermediate position between pure parsing benchmarks and downstream task benchmarks. It remains a structural extraction benchmark rather than a direct QA benchmark, but it explicitly chooses an output format—semantically rich Markdown—that is intended to be chunked, indexed, and used for RAG and knowledge construction (Li et al., 2024). This suggests a direct connection to later work showing that preprocessing quality materially affects downstream question answering in document-RAG pipelines (Santos et al., 30 Mar 2026).

2. Dataset design and source collections

READoc contains 2,233 documents drawn from two sources: READoc-arXiv and READoc-GitHub (Li et al., 2024). The two subsets were designed to expose different failure modes rather than to represent a single document genre.

Subset Size Characteristics
READoc-arXiv 1,009 documents arXiv preprints, 1996–2024, average pages 11.67, average token length 10,209.50, average heading depth 3.10
READoc-GitHub 1,224 documents GitHub README-style documents, 2008–2024, average pages 6.54, average token length 1,978.10, average heading depth 3.11

READoc-arXiv was extracted from arXiv preprints spanning 1996–2024, with six document types and eight disciplines (Li et al., 2024). This subset is challenging because it includes formulas, tables, multi-column layouts, and hierarchical academic structure (Li et al., 2024). READoc-GitHub consists of 1,224 GitHub README-style documents spanning 2008–2024 and covering 2,805 topics (Li et al., 2024). Its layouts are structurally simpler, but the paper emphasizes that heading inconsistency and ToC reconstruction remain difficult (Li et al., 2024).

The construction procedures reflect the benchmark’s attempt to produce high-quality reference Markdown rather than raw OCR targets. For arXiv documents, LaTeX sources were converted from LaTeX to HTML using LaTeXML, then from HTML to Markdown using a modified Nougat-style process (Li et al., 2024). For GitHub documents, README files were filtered by quality constraints such as being English, having more than 500 stars, and containing no HTML syntax; they were then converted from Markdown to PDF using Pandoc plus Eisvogel, and only documents with successful conversion and no errors or warnings were retained (Li et al., 2024).

This dual-source design is important because it prevents the benchmark from being reducible to a single document family. Academic PDFs stress formulas, dense tables, and multi-column scientific structure, whereas README-style documents expose failures in heading normalization, hierarchy recovery, and robustness to heterogeneous authoring conventions (Li et al., 2024). The later behavior of expert and generalist systems on the two subsets reinforces this distinction: Nougat-base is strongest on READoc-arXiv but degrades on READoc-GitHub, while Marker and GPT-4o-mini are comparatively stronger on GitHub-style documents (Li et al., 2024).

3. Task specification and the DSE Evaluation S3^3uite

READoc introduces the DSE Evaluation S3^3uite, which consists of three sequential modules: Standardization, Segmentation, and Scoring (Li et al., 2024). This evaluation design is central to the benchmark because systems being compared may emit different Markdown dialects or structurally similar but syntactically different outputs.

The Standardization module normalizes differences in formula boundaries, heading styles, table formats, and links or images (Li et al., 2024). Isolated formulas are normalized from variants such as %%%%2%%%% ... to ...... (Li et al., 2024). Headings are standardized to lines beginning with repeated #, and external URLs and image references are removed (Li et al., 2024). Markdown tables are converted to LaTeX tables for evaluation because LaTeX better represents complex row and column spans (Li et al., 2024). The purpose of this stage is to prevent superficial formatting variation from dominating the score.

The Segmentation module splits both prediction and reference into four semantic categories: headings, formulas, tables, and plain text (Li et al., 2024). Headings include all heading levels and support evaluation of hierarchical document structure; formulas include both embedded and isolated expressions; plain text is the residual textual content and includes basic formatting such as bold, italic, and lists (Li et al., 2024). Segmentation is implemented using regular expressions (Li et al., 2024).

The Scoring module evaluates both semantic units and reading order (Li et al., 2024). Semantic unit evaluation applies Edit Distance Similarity and vocabulary-level F1 to concatenated plain text, EDS plus TEDS to headings and ToC trees, EDS to embedded and isolated formulas, and EDS plus matched TEDS to tables after maximum bipartite matching (Li et al., 2024). Reading order detection is measured at two granularities: block level, where ordered lists of blocks are compared using Kendall’s Tau Distance Similarity, and token level, where sequential token lists are compared with the same metric (Li et al., 2024).

The explicit metric definitions are given in the appendix:

EDS(A,B)=1−ED(A,B)max⁡(∣A∣,∣B∣)EDS(A, B) = 1 - \frac{ED(A, B)}{\max(|A|, |B|)}

TEDS(T1,T2)=1−TED(T1,T2)max⁡(∣T1∣,∣T2∣)TEDS(T_1, T_2) = 1 - \frac{TED(T_1, T_2)}{\max(|T_1|, |T_2|)}

KTDS(X,Y)=1−2⋅Kdn(n−1)KTDS(X, Y) = 1 - \frac{2 \cdot K_d}{n(n-1)}

These choices make clear that READoc is evaluating multiple structural invariants simultaneously: lexical fidelity, hierarchical structure, table topology, and linearized order (Li et al., 2024). A plausible implication is that READoc’s evaluation philosophy treats layout understanding as indirectly observable through semantically normalized outputs rather than only through bounding-box detection.

4. Evaluated system classes and empirical findings

READoc evaluates four categories of systems under a unified setup: baselines, pipeline tools, expert visual models, and general VLMs (Li et al., 2024). The baselines are PyMuPDF4LLM, a bytecode parsing engine for digital PDFs, and Tesseract OCR with basic page segmentation (Li et al., 2024). Pipeline tools include MinerU, Pix2Text, and Marker (Li et al., 2024). Expert visual models are Nougat-small and Nougat-base (Li et al., 2024). General VLMs include DeepSeek-VL-7B-Chat, MiniCPM-Llama3-V2.5, LLaVa-1.6-Vicuna-13B, InternVL-Chat-V1.5, and GPT-4o-mini (Li et al., 2024).

On READoc-arXiv, the reported average scores are 40.55 for PyMuPDF4LLM, 35.13 for Tesseract OCR, 60.17 for MinerU, 64.39 for Pix2Text, 63.57 for Marker, 81.38 for Nougat-small, 81.42 for Nougat-base, and 57.98 for GPT-4o-mini (Li et al., 2024). On READoc-GitHub, Marker reaches 80.77, GPT-4o-mini 79.50, Nougat-small 74.86, and Nougat-base 74.12 (Li et al., 2024). These results establish that no single system family dominates all document types. Expert academic document models perform best on arXiv-style content, but their performance drops on GitHub-style documents, showing weak cross-domain generalization (Li et al., 2024).

The paper identifies several capability gaps. Reading order is comparatively easy, with systems often achieving high KTDS, especially at token level (Li et al., 2024). By contrast, ToC and heading hierarchy remain difficult because single-page systems struggle to recover document-wide structure and heading depth (Li et al., 2024). Tables and formulas are also persistent weak points, especially for general VLMs (Li et al., 2024). Performance declines on multi-column documents and on structurally deeper documents, which supports the claim that semantic unit evaluation is also measuring layout-analysis capability indirectly (Li et al., 2024).

READoc additionally examines a multi-page GPT-4o-mini setup and finds a trade-off: multi-page processing improves ToC construction but worsens table and formula conversion (Li et al., 2024). This suggests that broader context helps global structure while harming local precision. The same global-versus-local tension appears in later document systems. DocsRay, a zero-shot pseudo-TOC-guided hierarchical RAG system, uses hierarchical retrieval to improve long-document reasoning efficiency, but the paper notes that layout-heavy tasks requiring exact spatial relationships remain outside its core strength (Jeong et al., 31 Jul 2025). That comparison reinforces that document intelligence often decomposes into global structural reasoning and fine-grained local parsing, and that the two are not trivially optimized together.

5. Relation to surrounding benchmarks and document-reading research

READoc is best understood as part of an evolving benchmark lineage rather than as an isolated artifact. Earlier task-specific benchmarks had already established the importance of reading order, layout, or OCR individually. For example, LayoutReader introduced ReadingBank, a 500,000-page benchmark for reading order detection constructed automatically from WORD/XML metadata, and showed that layout-aware seq2seq modeling can nearly solve reading order detection on that benchmark (Wang et al., 2021). READoc extends beyond this scope by embedding reading order into a broader end-to-end structural extraction objective (Li et al., 2024).

Later benchmarks exposed limitations that READoc had already pointed toward. Dr. DocBench, introduced as a difficulty-aware benchmark for expert-level document parsing, uses parser-failure-based sampling from long multilingual books and evaluates page and block structure, reading order, and domain-specific visual content such as formulas, chemical structures, and music notation (Yang et al., 31 May 2026). Its findings—that strong performance on standard OCR benchmarks does not transfer reliably, that reading order degrades with longer windows, and that music transcription remains essentially unsolved—sharpen the claim that realistic document intelligence requires more than page-local transcription (Yang et al., 31 May 2026). In that sense, Dr. DocBench can be read as intensifying READoc’s realism criterion by shifting attention toward the hard long tail.

RealDocBench moves in another direction by replacing markdown-level similarity with field-level QA and layout evaluation on messy, regulated documents (Joshi et al., 5 Jun 2026). Its QA track contains 1,356 field-level questions over 581 documents, and its layout track contains 1,500 human-verified page images under a nine-class public taxonomy (Joshi et al., 5 Jun 2026). The benchmark reports both per-field and strict per-question accuracy, making explicit the downstream consequences of structural extraction errors (Joshi et al., 5 Jun 2026). This suggests that READoc’s PDF-to-Markdown objective is highly valuable for general-purpose evaluation, but that downstream production settings may require stricter task-grounded metrics than semantic reconstruction alone.

The downstream importance of READoc-style extraction is further supported by work on PDF-to-RAG pipelines. A systematic comparison of document conversion frameworks for domain-specific QA found that data preparation quality dominates downstream RAG performance, with metadata enrichment and hierarchy-aware chunking contributing more to QA accuracy than the conversion framework choice alone (Santos et al., 30 Mar 2026). Since READoc explicitly uses semantically rich Markdown as its output target, this later result suggests that benchmark progress on READoc is likely to matter directly for retrieval and question-answering systems built on parsed documents.

6. Limitations, implications, and continuing relevance

READoc’s limitations are largely implied by its design choices rather than stated as a separate limitations section (Li et al., 2024). Its coverage is bounded by two data sources, arXiv and GitHub, and therefore excludes many document genres that later benchmarks target more directly, such as regulated forms, multilingual books, or heavily distorted scanned pages (Li et al., 2024). Its target representation, Markdown, is strong but imperfect: it captures many structures, but not all document semantics (Li et al., 2024). Its evaluation also depends on normalization and regex-based segmentation, which are practical but may miss edge cases (Li et al., 2024). Finally, some systems may be disadvantaged by the full-document formulation if they were designed primarily for single-page processing (Li et al., 2024).

Despite these constraints, READoc remains significant because it shifted evaluation philosophy from isolated subtasks to unified document reconstruction (Li et al., 2024). It also clarified a structural fact that subsequent work has repeatedly confirmed: document intelligence is not reducible to OCR accuracy. Long-document continuity, hierarchy, tables, formulas, reading order, and downstream usability all require explicit consideration (Li et al., 2024). Later work on downstream QA, expert-domain parsing, and real-world regulated documents has not displaced this view; rather, it has specialized and extended it (Santos et al., 30 Mar 2026, Yang et al., 31 May 2026, Joshi et al., 5 Jun 2026).

A plausible implication is that READoc’s enduring contribution lies less in any one leaderboard and more in its task definition. By treating DSE as conversion from unstructured PDFs to semantically rich Markdown, it established a common interface for comparing bytecode parsers, OCR pipelines, expert visual models, and general VLMs under one benchmark (Li et al., 2024). That interface remains relevant precisely because later systems continue to be evaluated by how well they support realistic document reading, structured retrieval, and downstream reasoning rather than by page-local transcription alone.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to READoc.