Disentangle response variability from rater variability

Determine whether the lower rating consistency observed for comparative scientific analysis questions is caused by variation in AquiLLM's generated responses, differences in participants' scoring of ambiguous synthesis attempts, or a combination of both.

Background

The evaluation design required each participant to generate and rate their own responses, and no two participants rated the same generated response text. Consequently, variation in ratings for a given question combines genuine generation-to-generation differences with possible differences in how strictly participants applied the faithfulness scale.

Comparative Scientific Analysis questions exhibited both the lowest mean faithfulness and the lowest rating consistency. The paper explicitly states that the available data cannot determine whether this pattern reflects instability in AquiLLM's outputs, disagreement among evaluators about ambiguous synthesis responses, or both. Resolving these sources would require a study in which multiple raters independently assess the same fixed responses, potentially alongside repeated-generation experiments.

References

Because response variation and rater variation are completely confounded in this dataset, as described above, we cannot attribute this pattern to either source specifically: it may reflect AquiLLM's own output being less stable for comparative-analysis questions from one independent attempt to the next, participants disagreeing more about how to score genuinely ambiguous synthesis attempts, or some mixture of both, and this analysis provides no way to distinguish between these possibilities.

AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research  (2609.16519 - Boscoe et al., 15 Sep 2026) in Section 4.2, “Rating Consistency Across Independent Attempts” (Section \ref{sec:agreement})

We plan to rerun this faithfulness study, using the same query categories and evaluation methodology (Section~\ref{sec:querydesign}), against the updated system to test directly whether these changes reduce the synthesis and comparative-analysis failure modes identified in Section~\ref{sec:failuremodes}.

AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research  (2609.16519 - Boscoe et al., 15 Sep 2026) in Section 7, “Future Work” (Section \ref{sec:futurework})