Auditing and calibrating LLM-driven verification gates
Develop auditing protocols to calibrate large-language-model-driven quality gates used in verification and validation, incorporating adversarial probing, measurement of inter-agent disagreement, and evaluation against held-out physical measurements.
References
Finally, the two annotators are project collaborators rather than independent third parties, and the same pair supplies the judge calibration, so their labels may carry correlated expectation bias. An independent-annotator subset is left to future work.
Nine open questions will determine whether instrumented data matures into a recognised substrate for scientific machine learning. Verification of the verifier. Quality gates are LLM-driven; auditing their calibration requires adversarial probing, inter-agent disagreement, and held-out physical measurements.
Evaluation of physical faithfulness is itself an open problem: GenExam's MLLM judge and our binary-checklist evaluator both rely on a vision-LLM to answer the checks, and so remain proxies for expert judgment.
The live arm uses a local service, twelve contracts, and one writer; larger models, distributed services, correlated human decisions, adversarial faults, and longitudinal source drift remain open evaluations.