- The paper introduces a pipeline that deconstructs tax judgments into autonomous legal issues using an IRAC-based XML schema.
- It leverages a cascaded LLM extraction strategy with rigorous preprocessing and post-verification to ensure citation fidelity.
- Validation with expert annotations shows high precision, reducing citation hallucinations from 11.7% to 0.9%.
Introduction
The study introduces an automated, scalable pipeline for the structured extraction of legal reasoning from court judgments, specifically targeting the vast corpus (over 330,000 documents) of Italian tax-court decisions. The pipeline decomposes each judgment into individual, autonomous legal issues and represents each in a machine- and human-interpretable XML schema grounded in the IRAC (Issue, Rule, Application, Conclusion) framework and the logical legal syllogism. A distinctive feature of this approach is the integration of a citation hallucination-detection stage, specifically focused on curbing unreliability in LLM citation generation by verifying all extracted references against the underlying judgment text, normalized to standard identifiers.
Judicial systems output decisions in volumes that preclude manual reading or annotation. Structured representations at the issue level are critical for applications such as issue-based legal search, citation-network analysis, empirical studies of precedent flows, and the creation of benchmarking datasets for legal NLP. Prior work on legal NLP has focused primarily on attribution extraction, segmentation, summarization, and rhetorical role labeling, often using proprietary large models, few-shot prompt engineering, or less scalable schemas. Existing methods typically lack explicit citation faithfulness mechanisms, process at the judgment rather than issue level, or are not tailored for heterogeneous, mid-tier court corpora, such as Italian tax law, marked by high document entropy and sectioning irregularities.
Extraction Schema and Theoretical Foundations
The central design decision is to treat the "legal issue" as the atomic unit of extraction, reflecting modern jurisprudential theory and formal argumentation. Each issue is rendered as an XML block, encoding:
- Issue statement: A stand-alone whether-clause.
- Outcome: Direction (taxpayer, authority, other) and justification.
- Facts: Structured enumeration of relevant factual premises.
- Legal references: Exhaustive list, tagged by type (statute, case law, principle, administrative practice) and mapped to standard legal IDs (URN-NIR, CELEX, ECLI).
- Citation reason: Explanations corresponding to deeply discussed references.
- Judicial reasoning: Narrative explaining logical links from rule to outcome.
- Summary: Concise standalone abstractive overview.
This schema balances structure and expressiveness to optimize both automated extractability and downstream utility for human and machine consumers. It explicitly avoids over-fragmentation of reasoning chains (eschewing recursive syllogistic decomposition) in favor of robust, reproducible outputs, anchoring in the IRAC tradition but augmenting with citation provenance and type annotations.
Pipeline Architecture
Extraction proceeds in three major phases:
- Preprocessing: PDF-to-text conversion, section header identification, and exclusion of trivial/moot or ultra-short documents to focus only on judgments with substantive legal content.
- LLM Extraction: The DeepSeek V3 model is used for main extraction. Shorter judgments are processed with a single-prompt zero-shot approach; longer documents are handled by a cascaded two-step prompting strategy to decompose and assemble issues efficiently. Post-processing includes output format validation and syntactic repair to guarantee XML conformance.
- Citation Hallucination Control: After LLM extraction, all legal references in the output are cross-verified against standardized references automatically parsed from the judgment (using the Linkoln library), based on both text match and normalized identifiers. Items failing verification are purged.

Figure 1: Stylized visualization of the extraction and verification pipeline.

Figure 2: Flowchart of the hallucination check for items in <legal_references>, detailing the multi-step validation logic to ensure fidelity.
Validation Protocol and Empirical Results
A validation corpus of 50 judgments was doubly annotated by two tax law PhDs under double-blind conditions with detailed guidelines. Validation metrics include inter-annotator agreement (IAA), LLM-human agreement on issue extraction and legal citations, citation reason faithfulness, and qualitative Likert scoring of abstractive fields (completeness, correctness, form, satisfaction).
Salient findings:
- Issue extraction: IAA of 88.1% (F1), LLM-vs-human F1 between 80.3% and 87.9%, with error profile dominated by under-extraction, not hallucination. LLM precision exceeded 93%.
- Citation extraction: IAA of 97.1% (F1), LLM-vs-human F1 of ~74%, indicating substantial alignment on legal reference selection given the heterogeneity of source material.
- Citation hallucinations: Pre-filter hallucination rate of 11.7% overall, soaring to 69% for "principle"-type references. Post-filter, only 0.9% hallucination remained, at the minimal expense (3%) of valid references. The hallucination filter thus functions as a high-specificity, high-precision binary classifier for reference fidelity.
- Qualitative fields: All main fields (fact premises, reasoning, issue outcome, summary) average above 4.6/5 with high (≥93%) IAA on Likert scales, supporting robust interpretability for both experts and downstream machine tasks.
Implications and Future Directions
This work establishes a robust, scalable paradigm for issue-level legal information extraction that is appropriate for domains with highly variable document formality and structure. The approach significantly advances citation-level faithfulness and provides a framework for understanding the empirical trade-off between extractive granularity and LLM reliability. By situating the extraction at the legal issue layer, the method supports future graph-based citation analysis, issue-based retrieval, and the construction of better benchmarks for future legal NLP models.
Key practical and theoretical implications include:
- Quantification and mitigation of LLM hallucination in domain-critical settings.
- Empirically grounded tradeoff analysis for cost-effective application of foundation models to law.
- A basis for cross-jurisdictional adaptation by converting domain-specific citation forms to universal legal resource identifiers.
- Enabling construction of higher-fidelity legal citation networks, improved legal search at the issue level, and large annotated datasets for legal reasoning research.
Conclusion
The presented pipeline achieves near-expert-validated fidelity in converting raw court judgments to structured legal issues, with strong hallucination resistance and excellent qualitative evaluations on abstractive fields. The methodology generalizes readily to other jurisdictions and corpora and provides a practical blueprint for high-throughput, trustworthy legal information extraction at the scale of modern judicial systems.
Future trajectories should focus on schema adaptation to other legal subdomains, porting to multilingual and multi-jurisdictional corpora, enhanced retrieval techniques leveraging the constructed citation networks, and benchmarking model variants in cost–performance space as LLMs improve. Replication on larger and more heterogeneous validation sets is warranted to further refine generalizability and push toward fully automated legal reasoning pipelines.
Reference:
"From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control" (2607.03325)