Disentangle response variability from rater variability
Determine whether the lower rating consistency observed for comparative scientific analysis questions is caused by variation in AquiLLM's generated responses, differences in participants' scoring of ambiguous synthesis attempts, or a combination of both.
References
Because response variation and rater variation are completely confounded in this dataset, as described above, we cannot attribute this pattern to either source specifically: it may reflect AquiLLM's own output being less stable for comparative-analysis questions from one independent attempt to the next, participants disagreeing more about how to score genuinely ambiguous synthesis attempts, or some mixture of both, and this analysis provides no way to distinguish between these possibilities.
We plan to rerun this faithfulness study, using the same query categories and evaluation methodology (Section~\ref{sec:querydesign}), against the updated system to test directly whether these changes reduce the synthesis and comparative-analysis failure modes identified in Section~\ref{sec:failuremodes}.