Equivalence of structured and original prompt formats
Establish whether the structured answer-field rendering is equivalent to the original candidate-aggregation prompt for purposes of measuring accuracy and causal answer-field effects.
References
On the original aid materials the same pipeline reproduces the direction of both published effects while neither interval clears zero, which points at the materials without saying what in them carries the phenomenon. We offer that as a conjecture, not a finding.
— The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits
(2609.09048 - Vohra et al., 8 Sep 2026) in Discussion and Conclusion
Equivalence with the original prompt was not established.
— Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation
(2608.18379 - Farmanfarmaian, 18 Aug 2026) in Section 3.3, “A structured answer-field intervention”; Appendix F, “The structured answer-field intervention: details”
A full-cohort rescoring under a single symmetric extraction rule has not been carried out and remains the outstanding measurement; the convention correction, by contrast, has been applied to all 42 deployments.
— HEPToolBench 1.2: Testing How Reliably Language Models Can Drive Particle Physics Software
(2608.28232 - Singh et al., 28 Aug 2026) in Section 2.2, paragraph “Convention and extraction audit” (Sec. scoring-policy)