Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act

Published 16 Jun 2026 in cs.CY, cs.AI, and cs.CL | (2606.18158v1)

Abstract: LLMs now produce legal text of at least median quality, yet no existing benchmark can evaluate whether they perform doctrinal legal reasoning, which forms the interpretive core of legal work, rather than the ancillary, paralegal tasks that most current legal-AI evaluations measure. This measurement gap is not only methodological but legal: the EU AI Act makes "appropriate accuracy" a binding requirement for high-risk AI used in the judicial domain, yet that requirement cannot acquire operational content without the very doctrinal-reasoning benchmark the field lacks.

Authors (1)

Summary

  • The paper identifies a critical gap in doctrinal legal reasoning benchmarks required by the EU AI Act and introduces a taxonomy of 21 failure modes.
  • The paper critiques existing benchmarks for legal AI, highlighting their inability to capture the internalism, normativity, contestability, and coherence of EU doctrinal reasoning.
  • The paper underscores the regulatory necessity for robust doctrinal benchmarks to ensure the accuracy and compliance of high-risk judicial AI systems in the EU.

Overview and Motivation

This paper identifies and analyzes a critical methodological and regulatory gap in the measurement of LLM performance on doctrinal legal reasoning within the EU context, specifically concerning the mandates of the EU AI Act. While current benchmarks for "legal AI" overwhelmingly assess paralegal tasks or exam-style question answering, they fail to capture the core activity of legal practice: doctrinal legal reasoning as it surfaces in academic, judicial, and practitioner work.

The disconnect is not merely academic. The EU AI Act imposes binding accuracy requirements for high-risk AI used in the judicial domain (Article 15), yet the absence of benchmarks for doctrinal legal reasoning renders this obligation void of operational content. The paper makes two diagnostic claims: (1) current benchmarks are systematically unsuited for doctrinal legal reasoning, and (2) the AI Act’s regulatory framework not only exposes but also compels remedying this gap. The principal constructive contribution is a taxonomy of failure modes that any rigorous doctrinal reasoning benchmark in EU law must be able to detect.

The author distinguishes doctrinal legal reasoning from doctrinal legal research, focusing on the cognitive process by which legal norms are interpreted, synthesized, and situated within the broader legal system. Four core features that a benchmark must capture are systematically articulated:

  • Internalism: Reasoning is situated within the legal order, relying on legal sources and the legal system’s own structures and standards of correctness. The primacy and hierarchy of sources are paramount.
  • Normativity: Interpretation is inherently normative, not merely descriptive; it involves evaluative judgments about what the law requires in its best reading, independently of prevailing sociopolitical views or mere statistical majority.
  • Contestability: Legal rules are structurally open-textured, making multiple defensible interpretations admissible. Doctrinal reasoning anticipates and addresses plausible contestation.
  • Coherence: Sound doctrinal analysis integrates individual rules and decisions into a systematically ordered and principled account, evaluating candidate readings by their fit with underlying principles and the legal system’s architecture.

The paper argues that these features are present and even intensified in EU law due to supranational structure, doctrinal autonomy, multilingualism, layered sources, and a deep commitment to legal coherence.

The Regulatory Imperative and Measurement Problem

The author systematically demonstrates that the regulatory architecture of the EU AI Act, specifically Article 15, renders the benchmark gap legally consequential. The "appropriate accuracy" requirement for high-risk judicial AI systems cannot be fulfilled—or even meaningfully defined—without doctrinal reasoning benchmarks. Furthermore, Article 6(3) creates indirect regulatory pressure for such benchmarks by allowing providers to claim that their systems do not materially influence judicial decisions, a claim that in practice demands standardized empirical substantiation.

The paper incisively critiques the limits of existing benchmarks:

  • ECtHR-CASES and LexGLUE: Focused on outcome prediction and classification tasks, neglecting inner structure, contestability, and the systematization that mark doctrinal reasoning.
  • LegalBench and LEXam: While tasks move in the direction of free-form textual analysis, they remain grounded in narrow jurisdictions, lack systemic coherence checks, and default to answer keys that flatten contestability.
  • BenGER: Engages with the internal structure of German law and incorporates contestability as distributional expert assessment, yet fails to address the complex system integration and cross-norm coherence the EU context requires.
  • Harvey LAB: Prioritizes practical work completion over doctrinal soundness.

The author’s survey reveals no existing benchmark that appropriately operationalizes doctrinal legal reasoning in the EU context.

The Failure Taxonomy: Towards Doctrinal Soundness in Evaluation

The central constructive contribution is a taxonomy of 21 failure types that any credible doctrinal reasoning benchmark must be able to detect. This taxonomy is informed by both the structure of EU law and the methodological literature on legal reasoning.

Key classes include:

  • Failures in Source and Authority Recognition: Confusion of norm hierarchy, improper weighting of judicial and soft law materials, and loss of autonomous meaning in EU legal terms.
  • Failures in Doctrinal Doctrines and Effects: Inability to handle doctrines like supremacy, direct effect, consistent interpretation, and procedural nuances (e.g., preliminary reference binding effects).
  • Interpretive and Methodological Failures: Treating frequency as correctness, inability to adapt interpretive method to context, and improper use of multilingualism and recitals.
  • Evolutive and Contestability Failures: Blindness to legal evolution, omission of ongoing contestation or uncertainty, and inability to indicate boundary cases or acte clair.
  • Coherence and Integration Failures: Producing analyses that are locally correct but systemically incoherent, ignoring cross-instrument and cross-norm consistency, inflating the authority of non-binding sources, and mismatched citation-proposition relationships.

Several failure modes align with challenges recently identified in LLMs, such as the degradation of performance in long-context document integration (Nagl et al., 27 May 2026).

The taxonomy directly informs not only technical benchmark design but also the construction of input-output pairs, gold standard references (or expert panels for contestable cases), and scoring criteria. It recognizes that doctrinal competence cannot be reduced to statistical text pattern matching or completion, but requires structured, legally meaningful argumentation, and system-wide integration.

Implications and Outlook

The absence of doctrinal reasoning benchmarks for EU law is no longer an academic lacuna but a regulatory liability, given the operational and compliance requirements of the AI Act. Theoretical advances in LLMs and retrieval-augmented generation cannot be assessed in legal domains without reference to meaningful doctrinal benchmarks. The taxonomy offered here is essential for any interdisciplinary collaboration aimed at building such instruments.

On a practical level, the development of doctrinal benchmarks will shape the industrial policy and technological trajectory of legal AI within the EU. Model capabilities will concentrate where measurement is feasible; thus, continued reliance on US-centric or paralegal benchmarks will skew innovation away from core legal reasoning in EU contexts. The taxonomic framework opens a path for the legal discipline to concretely engage and guide technical work.

Conclusion

The paper establishes the necessity—methodologically and legally—of benchmarks capturing doctrinal legal reasoning for the EU AI Act’s effective implementation regarding high-risk judicial AI systems. It exposes the unsuitability of current benchmarks and provides a detailed taxonomy of failure cases that should underpin future benchmark construction. The research delineates the collaborative task ahead for computer scientists and legal scholars and positions doctrinal benchmarking as a core instrument for both regulatory compliance and trustworthy AI innovation in the legal domain.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.