Auditing and calibrating LLM-driven verification gates

Develop auditing protocols to calibrate large-language-model-driven quality gates used in verification and validation, incorporating adversarial probing, measurement of inter-agent disagreement, and evaluation against held-out physical measurements.

Background

The instrumentation pipelines rely on LLM-driven gates for quality control and verification steps. Ensuring these gates are reliable requires independent calibration and robust auditing procedures.

The authors highlight the need for adversarial testing and disagreement-based diagnostics, anchored by comparison to physical measurements, to validate and monitor these gate mechanisms.

References

Finally, the two annotators are project collaborators rather than independent third parties, and the same pair supplies the judge calibration, so their labels may carry correlated expectation bias. An independent-annotator subset is left to future work.

GANDR: Claim Auditing for Verifiable Legal Answer Generation  (2609.10293 - Qian et al., 9 Sep 2026) in Appendix A, Section “Human validation of atomic cite_verify”

Nine open questions will determine whether instrumented data matures into a recognised substrate for scientific machine learning. Verification of the verifier. Quality gates are LLM-driven; auditing their calibration requires adversarial probing, inter-agent disagreement, and held-out physical measurements.

Instrumented data for causal scientific machine learning  (2606.07865 - Wilke, 5 Jun 2026) in Section 7, Methodological questions for the community, Item 3

Evaluation of physical faithfulness is itself an open problem: GenExam's MLLM judge and our binary-checklist evaluator both rely on a vision-LLM to answer the checks, and so remain proxies for expert judgment.

Towards Physics-Faithful Generation of Scientific Diagrams  (2608.13112 - Zhang et al., 13 Aug 2026) in Section 4.1, “Future Work” (Discussion)

The live arm uses a local service, twelve contracts, and one writer; larger models, distributed services, correlated human decisions, adversarial faults, and longitudinal source drift remain open evaluations.

Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures  (2609.10969 - Zheng et al., 10 Sep 2026) in Section 'Engineering Implications and Validity', subsection 'External validity'