Calibrate per-cell power simulations

Calibrate a per-cell power simulation for the negative-binomial generalized linear mixed model that matches the single-prompt, single-model audit design and yields iteration-count recommendations that quantitatively align with the generalizability-theory D-study.

Background

The paper’s primary simulation pools eight prompts and three models, causing statistical power to saturate at very small iteration counts because each simulated test contains many observations. The authors therefore caution that these results are an upper bound and that the generalizability-theory D-study governs recommendations for individual prompt–model cells.

A simulation matching the actual binding design—one prompt, one model, and repeated iterations—would provide directly applicable power estimates and could reconcile the pooled-prompt power results with the D-study-based iteration thresholds.

References

A calibrated per-prompt simulation that matches the single-cell audit design (one model, one prompt, $n$ iterations) is retained as future work; the current table should be read as an upper bound on power under pooled-prompt analysis, and the D-study as the binding constraint for per-cell audits.

The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations  (2609.04047 - Żatuchin, 3 Sep 2026) in Section 5.1, subsection “GLMM Simulation-Based Power (RQ1)”