Assess inter-rater reliability and elicitation sensitivity

Assess the inter-rater reliability of the structured gain-function elicitation procedure and determine how sensitive the optimised graphical multiple testing procedure is to variability in the elicited stakeholder scores.

Background

The paper proposes eliciting the relative value of incremental label claims from a cross-functional panel of clinical, commercial, and regulatory stakeholders. Individual ratings are aggregated, typically using the median of final-round scores, and the resulting gain function is then used to optimise the graphical multiple testing procedure.

The authors note that the additive gain model simplifies elicitation but may not capture diminishing returns or other dependencies among label claims. They explicitly identify two unresolved methodological issues: the reliability of agreement among raters and the extent to which differences or uncertainty in elicited scores alter the graph selected by the optimisation.

References

The inter-rater reliability of the procedure and the sensitivity of the optimised graph to variability in elicited scores remain areas for formal assessment.

— Gain-function optimisation of graphical multiple testing procedures for confirmatory clinical trials  (2609.19994 - Spiers et al., 17 Sep 2026) in Section Discussion, paragraph beginning “For practical elicitation”