Determine whether observer opinions or authority impersonation are more effective attacks

Determine whether appending an unverified observer opinion to the state of a typed decision request causes more decision flips than injecting an authority-impersonation command into a question string.

Background

The benchmark finds that an observer opinion appended to the state flips 12.1% of decisions, whereas authority impersonation in a question string flips 10.1%. Although the observer opinion has the higher point estimate, the confidence interval for their difference includes zero, so the study does not establish which attack is genuinely more effective. The authors therefore treat the two attacks as statistically tied.

References

The difference between the two strongest attacks is not resolved, so the paper calls them tied.

— JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models  (2609.31142 - Hu et al., 25 Sep 2026) in Appendix E, paragraph “Observer opinion and authority impersonation are statistically tied”