- The paper introduces BenGER, a three-corpus benchmark of 1,127 German legal reasoning tasks evaluated across 12 LLM systems using the Obersatz–Definition–Subsumtion–Ergebnis framework.
- The paper finds that closed flagship systems lead performance, with top models achieving 83.8–85.1 raw points on Benchathon and lexical metrics such as METEOR correlating more strongly with rubric scores than embedding metrics.
- The paper shows that LLM judging can approximate human-reviewer agreement, while AI-assisted human solutions scored 19.9 raw points higher than traditional solutions, though sampling bias and judge-style alignment limit causal conclusions.
BenGER is a benchmark for evaluating LLM systems on subsumption-based legal reasoning in German law, released alongside an evaluation methodology that treats legal grading as a distribution over plausible expert judgments rather than a single ground-truth label. The benchmark comprises three corpora — 581 publicly available exam-style cases from the Zeitschrift für das Juristische Studium (ZJS), 15 newly collected Benchathon tasks with 220 human solutions each under two working conditions, and 531 short doctrinal reasoning items — evaluated with 12 contemporary LLM systems (14,614 generations total), a rubric-aligned LLM-as-a-Judge, and a multi-rater human-grading layer used to validate the judge.
Motivation: subsumption and noisy human grading
The benchmark targets the Gutachtenstil scaffold — Obersatz (norm hypothesis), Definition, Subsumtion, Ergebnis — that structures German legal education and state examinations. Unlike classification or multiple-choice legal benchmarks, grading of a Falllösung rewards how each norm element is anchored in the specific fact pattern, which makes the task a demanding structured-generation problem and simultaneously amenable to rubric-based evaluation: each scaffold step maps to a dedicated rubric dimension. The authors motivate their evaluation philosophy with Hufeld's finding that identical German law exam answers receive grades varying by an average range of 6.47 points on the 0–18 scale, with only 42% of grades within ±1 point of the mean. This justifies treating evaluation as a distribution over expert assessments and validating any automated judge against inter-human variability rather than against a single label. The closest prior benchmark, LEXam, includes German exam questions but does not couple free-text items to the Obersatz–Definition–Subsumtion–Ergebnis rubric used in German grading.
Dataset design
The ZJS corpus (581 cases) spans civil, criminal, and public law from early undergraduate exercises through first state-examination material, each with an expert reference solution, and supports at-scale automated evaluation. The Benchathon corpus (15 intermediate-difficulty tasks) was solved in a controlled setting by 33 participants — predominantly law students, plus Referendare, graduates, and a small layperson cohort — under a two-hour limit, with each task randomly assigned to either a traditional condition (statutes, databases, literature allowed) or a co-creation condition (AI tools additionally allowed), yielding 65 traditional and 155 co-creation solutions. The Doctrinal Principles corpus provides 531 short question–answer pairs with Ja/Nein decision labels. The benchmark is explicitly evaluation-only, with no train/test split.
Experimental setup
Twelve systems were evaluated black-box through provider APIs: six closed-weight (OpenAI, Anthropic, Google) and six open-weight via a managed inference provider, with provider-recommended temperatures and token budgets and a uniform German-language prompt mirroring the Gutachtenstil scaffold. Scoring uses a ten-dimension rubric judge (GPT-5-mini as primary judge) covering result correctness, issue identification, legal grounding, doctrinal knowledge, subsumption quality, problem depth, methodological style, organization, terminology, and formal correctness, aggregated to a 100-point raw score and the German 0–18 grade scale; a pass requires grade ≥ 4 (raw ≥ 50). The judge is cross-validated on a 45-solution subset (one traditional, one co-creation, one LLM solution per task), each receiving three blind human reviews plus one author-informed creator review (180 reviews total).
Closed flagship systems lead every corpus. On Benchathon, Sonnet-4.6 scores 85.1 raw points, Opus-4.7 84.2, and GPT-5.4 83.8, all with 100% pass rates; on ZJS, GPT-5.4 leads at 84.9 raw (14.1 grade points, 99% pass) versus Llama-4 at 59.6 (6.8 grade points, 76% pass). On Doctrinal Principles, the same cluster leads, with Ja/Nein accuracy ranging from 64% to 77% (Gemini-3.1-Pro highest at 77%). A consistent finding is that closed-API dominance narrows sharply below the flagship tier: the flagship-vs-open-weight gap is 18 raw points on Benchathon but only 9 on ZJS (closed-vs-open: 16 to 7). Per-task variance is also tier-dependent — top systems show per-task standard deviations of 6–8 raw points versus 10–15 for mid-tier and smaller open-weight systems.
The human cohort scores in the mid-range of the LLM distribution: pooled human solutions average 76.6 raw points under judge scoring, with unaided traditional solutions at 63.3 and co-creation solutions at 82.0. The authors frame all tier comparisons as system-level rather than model-level effects, since black-box API access entangles the open–closed split with provider-side prompting, decoding, and possible routing.
Metric validity and co-creation effects
A counterintuitive result concerns automatic metrics: lexical metrics correlate with the rubric judge more strongly than embedding-based ones across all corpora. On Benchathon, METEOR (r = +0.61), ROUGE and BLEU (r = +0.56) lead, while BERTScore (r = +0.38) and sentence-embedding similarity (r = +0.28) trail. The authors attribute this to the highly templated structure of doctrinal German exam writing, in which reference and high-quality answers share substantial n-gram material around statutory citations and the subsumption scaffold. At the system level, METEOR's Spearman correlation with the judge ranking reaches +0.80, but lexical metrics would misrank the mid-tier, motivating the rubric judge as the primary instrument.
On human–AI co-creation, the judge rates co-creation solutions at 82.2 raw versus 62.3 for traditional solutions — a difference of +19.9 points (bootstrap 95% CI [+15.2, +24.6]) — with the co-creation mean falling inside the closed-flagship-tier CI (mean 83.3). Under blind human reviewers the difference is smaller (+13.6 points, CI [+0.5, +26.7]). The authors flag that the comparison rests on unbalanced, non-independent samples with a non-clustered bootstrap, and note a judge-style-alignment caveat: the LLM judge may favor LLM-flavored prose, which bears directly on the co-creation estimate.
Grading reliability of the LLM judge
The central reliability result is that the LLM judge can substitute for a human reviewer without degrading pool agreement. The judge agrees with the per-solution mean of blind reviewers at Pearson r = 0.78 (Spearman ρ = 0.73), MAE 14.7 raw points (≈3.5 grade points), and Cohen's κ = 0.67 on the pass/fail decision — while the blind reviewers themselves disagree by a within-solution max–min spread of 23.9 raw points (5.2 grade points), consistent with Hufeld's inter-human range. The Calderon-style alternative-annotator check is the sharpest test: on 30 matched annotations, replacing a blind human with the judge yields r = 0.96 correlation with the full-human pool, identical to the r = 0.96 obtained by simply dropping that human. The blind pool achieves ICC(2,k) = 0.84.
Judge self-consistency exceeds human consistency: k = 3 re-runs of GPT-5-mini yield a mean within-cell standard deviation of 2.51 raw points, far below the 23.9-point inter-human spread. Across three judge families, a strictness gradient appears — Opus-4.7 scores on average 21.5 raw points below GPT-5-mini — but pairwise rank agreement remains high (Spearman ρ between 0.82 and 0.92). Dimension-level agreement is highest on structured criteria (result correctness r = 0.81, issue identification r = 0.79) and lowest on presentation-oriented dimensions (organization and terminology r = 0.55, formal correctness r = 0.45), mirroring the human inter-rater pattern.
Error analysis identifies five recurring failure modes, including correct outcomes reached without fact-bound subsumption, missing subsumption in short answers, and — most diagnostic of judge behavior — five cases where the judge praises fact-bound subsumption (≥ 0.60 of max) yet the final answer is wrong, demonstrating that rubric scoring captures reasoning/outcome dissociation that single-number metrics cannot.
Limitations
The authors are explicit about the bounds of their claims. The full RQ1–RQ2 leaderboard rests on a single judge model (GPT-5-mini); because the primary judge shares a model family with the leaderboard-leading GPT-5.4 systems, same-provider preference cannot be fully ruled out, and the authors state the leaderboard should be re-scored with a non-OpenAI judge before strong cross-judge claims. The Benchathon validation subset (15 tasks, 45 graded picks) limits statistical power. The rubric and judge prompt were co-developed for this benchmark and validated only as a pair. ZJS material is publicly indexed, so pretraining contamination is possible; the agreement of ZJS ordering with the private corpora bounds catastrophic differential contamination but not proportional contamination. The Google efficiency-tier slot uses two different models across corpora, and the Doctrinal Principles corpus uses a different four-dimension rubric, precluding per-dimension cross-corpus comparison. Several analyses are reported descriptively without multiple-comparison correction.
Conclusion
BenGER provides the first benchmark coupling open-ended German subsumption-based legal reasoning with a validated evaluation protocol under noisy expert judgment. Its principal empirical findings are that closed flagship systems lead across all three corpora with a narrowing gap at the mid-tier, that lexical metrics outperform embedding metrics against rubric scoring on this templated register, that human–AI co-creation improves exam performance by roughly 20 raw points to flagship-level quality, and that an LLM judge substitutes for a human rater without degrading multi-rater pool agreement. The benchmark is released with corpora, rubric, judge prompts, and a fully reproducible analysis pipeline, positioned as one signal among several for legal-domain LLM evaluation rather than a standalone ground truth.