Papers
Topics
Authors
Recent
Search
2000 character limit reached

Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?

Published 21 Aug 2026 in cs.AI | (2608.21089v1)

Abstract: Integrating LLMs into the Indian judiciary promises access to justice but introduces severe risks. We identify the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized 'precedent overfitting' bias. Phase I of our socio-technical audit tested ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the shift toward statutory enforcement of specific performance. We introduce the High-Confidence Error Rate (HCER) to quantify incorrect verdicts delivered with dangerous certainty (>= 9 on a 1-10 scale). All models struggled with statutory updates. Meta AI proved most vulnerable (31.7% HCER), frequently misapplying pre-amendment rules with a 9.1/10 mean confidence, followed by Perplexity (15.0%) and ChatGPT (6.7%). Phase II investigated human vulnerability to this overconfidence via a survey of Indian law students (N=380). Verification often functions as a reactive adaptation to machine hallucinations: students encountering fabricated citations reported higher verification scores (4.2/5) than those with no such encounters (2.8/5). Furthermore, while 81.6% knew submitting hallucinated cases can lead to contempt-of-court, 71.1% received no formal training on ethical AI use. We propose shifting toward adversarial legal research pedagogy and implementing source-grounded verification architectures to prevent systemic professional negligence.

Summary

  • The paper identifies and quantifies the 'inertia of confidence' in AI, noting that models deliver incorrect legal verdicts with high-confidence scores (High-Confidence Error Rate—HCER) reaching up to 31.7% for Meta AI in contract law cases.
  • The study reveals that 71.1% of surveyed Indian law students have not received formal training on ethical AI use, correlating with student reaction relying on machine outputs despite contenders of court issues.
  • The authors propose interventions including adversarial legal research modules, a Verifiable Authority Index (VAI), and an Internal AI Verification Protocol (IAVP) to mitigate the detection of high-confidence errors and bridge the identified socio-technical risk gap.

This paper presents a dual-layered socio-technical audit of LLMs in the Indian legal domain, combining a black-box evaluation of three consumer-facing AI systems against a 60-case battery on Indian contract law with a cross-sectional survey of 380 Indian law students. The central contribution is the identification and quantification of what the authors term the "inertia of confidence"—a pattern in which models deliver incorrect legal verdicts with near-maximum self-reported confidence—and the introduction of the High-Confidence Error Rate (HCER), a metric capturing the proportion of incorrect outputs delivered at a confidence score of 9 or above on a 10-point scale. The paper links this algorithmic overconfidence to human overreliance among law students, arguing that generic human-in-the-loop safeguards are insufficient without adversarial pedagogy and source-grounded verification architectures.

Motivation and conceptual framing

The study is situated within India's accelerating judicial adoption of AI, including SUVAS for judgment translation and SUPACE for research assistance. The authors challenge two assumptions prevalent in prior work: first, that LLM metacognitive calibration is generally sound (as suggested by Kadavath et al.'s finding that models "mostly know what they know"), and second, that legal AI errors are attributable primarily to knowledge cutoffs or temporal lag. They propose an alternative mechanism, "precedent overfitting," in which the statistical volume of historical pre-amendment judgments disproportionately influences model outputs relative to recent statutory overrides—specifically the Specific Relief (Amendment) Act, 2018, which converted specific performance from discretionary equitable relief into a presumptive statutory mandate. Notably, the paper concedes that because its design is black-box, it cannot observe training weights, attention mechanisms, or retrieval rankings; precedent overfitting therefore remains a hypothesis consistent with observed error patterns rather than an established causal account.

Phase I methodology

The technical audit evaluated ChatGPT (GPT-5.2) via its web interface, Meta AI via WhatsApp, and Perplexity AI (Sonar) via its free-tier web interface during the first quarter of 2026. Selection was deliberately based on access modality rather than exhaustive model coverage: conversational web interface, zero-barrier messaging integration, and native retrieval-augmented generation (RAG), respectively. Gemini and Claude were excluded on these grounds.

The 60-case "judicial agent benchmark" comprises six categories of ten cases each: five doctrinal control sets covering offer/acceptance, capacity/consent, consideration/lawful object, discharge/frustration, and damages/enforcement under the Indian Contract Act, 1872, plus a temporal stress test of ten cases targeting the 2018 Specific Relief amendments. Scenarios were constructed from established case facts but anonymized—party names, dates, and geographic markers stripped—to reduce memorization-based retrieval. Each session was initialized with a standardized "Judicial Persona" prompt requiring a definitive verdict, supporting statutory authority, and a self-assessed confidence score from 1 to 10.

Two methodological caveats are acknowledged by the authors themselves. First, because the persona prompt demanded definitive conclusions from an apex-court perspective, observed confidence may partly reflect prompt-induced response pressure; HCER should be read as high-confidence error under this task framing rather than an intrinsic model trait. Second, HCER is a risk-oriented metric rather than a formal calibration estimate.

Phase I results

The results show a sharp dichotomy between historical doctrine and modern statutory shifts:

Case Category ChatGPT Perplexity Meta AI
Offer, Acceptance & Communication (1–10) 10/10 9/10 8/10
Capacity, Consent & Formation (11–20) 10/10 8/10 8/10
Consideration & Lawful Object (21–30) 9/10 8/10 7/10
Discharge, Frustration & Restitution (31–40) 9/10 8/10 7/10
Damages, Terms & Enforcement (41–50) 8/10 7/10 6/10
Specific Relief & 2018 Amendments (51–60) 7/10 6/10 5/10
Total accuracy 88.3% 76.7% 68.3%
HCER 6.7% 15.0% 31.7%

Accuracy degraded monotonically toward the modern-statutory category across all three systems, while mean self-reported confidence remained high (8.8/10 for Perplexity to 9.4/10 for ChatGPT). Meta AI's HCER of 31.7%—nearly one in three answers incorrect yet delivered at confidence ≥9—is the paper's most striking quantitative result, particularly given WhatsApp's ubiquity as a zero-barrier legal assistant in India. Qualitative failure analysis identified two modes: a prospective-versus-retrospective error in which Meta AI treated the temporal application of the 2018 Amendment as settled despite the recall of Katta Sujatha Reddy v. Siddamsetty Infra Projects (asserting 10/10 confidence), and statutory fabrication, where Perplexity hallucinated a non-existent "7-day cure period" in place of the mandatory 30-day notice under Section 20(2).

The audit also documented an unanticipated behavioral anomaly termed "instructional over-compliance": when Meta AI exhausted the 60-case input, it spontaneously generated 30 additional synthetic scenarios complete with fabricated statutory authorities and high confidence scores, mimicking formatting constraints without any grounding in actual Indian law. The authors attribute this to consumer-model optimization for politeness and conversational continuity, noting the operational risk of practitioners mistaking structurally fluent filler for binding jurisprudence.

Phase II: student survey

The survey of N=380N=380 LLB students used a purposive convenience sample approximating the Cochran baseline (n0384n_0 \approx 384), recruited across National Law Universities, private schools, and state-affiliated institutions. On hallucination exposure, 42.1% reported multiple encounters with fabricated citations, 36.8% occasional encounters, and only 21.1% none. Verification behavior differed sharply by exposure history: students reporting multiple hallucination encounters had a mean verification frequency of 4.2/5 versus 2.8/5 for those reporting none. The authors frame this as preliminary cross-sectional evidence for a "reactive verification response"—verification acquired through encounter with machine hallucinations rather than through training—while explicitly conceding that the cross-sectional design cannot establish causal ordering.

The institutional findings are stark: 81.6% of students were aware that submitting hallucinated cases can trigger contempt-of-court consequences, yet 71.1% reported receiving no formal training on ethical AI use. Mean job-displacement anxiety was 3.34/5, descriptively associated with the absence of institutional safeguards—a state the authors label "unprotected accountability."

Discussion and policy proposals

The synthesis argues for a "Socio-Technical Risk Gap" and a "Double Blindspot": algorithmic temporal lag (models defaulting to pre-amendment rules) compounded by pedagogical lag (students unequipped to audit black-box outputs). The 52.6% of students who only "sometimes" cross-verify are identified as vulnerable to cognitive offloading precisely when outputs are confident. Three interventions follow: (1) mandatory "Adversarial Legal Research" modules in the LLB curriculum, including hallucination-detection labs and red-teaming of AI outputs; (2) a Verifiable Authority Index (VAI) requiring programmatic linkage of model outputs to verified reporters such as SCC, AIR, or e-SCR; and (3) an Internal AI Verification Protocol (IAVP) providing structured source-grounded human audit layers in firms and courts.

Limitations and open questions

The paper is candid about several constraints. The black-box design precludes establishing precedent overfitting as a causal mechanism; only the correlation between error patterns and statutory recency is demonstrated. The persona prompt may inflate confidence scores, limiting generalization of HCER beyond this framing. The student sample is non-probabilistic, restricting inferential claims, and the verification-exposure association is descriptive rather than causal. Model coverage excludes Gemini and Claude, and Meta AI's underlying checkpoint was not exposed by the interface, complicating attribution. Open questions include whether professional experience among advocates and judges naturally mitigates the automation bias observed in students, and whether RAG architectures can be engineered to suppress precedent overfitting without sacrificing fluency.

Conclusion

The paper quantifies a domain-specific calibration breakdown in legal LLMs—HCER peaking at 31.7% for Meta AI—concentrated precisely where statutory change conflicts with voluminous historical precedent, and pairs it with survey evidence that student verification practices are largely reactive rather than trained. Its principal claim, that human-in-the-loop frameworks alone cannot close the resulting risk gap, rests on descriptive evidence but is supported by converging algorithmic and behavioral data. The proposed remedies—adversarial pedagogy, verifiable authority indices, and structured institutional verification protocols—constitute a concrete agenda whose effectiveness remains to be empirically tested.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.