Safety of prediction-based reward under manipulable reality

Establish how to design reality-settled scoring so that systems rewarded for predicting reality cannot instead manipulate the world, measured environment, settlement substrate, or claim-selection process to make outcomes easier to predict.

Background

The paper emphasizes that replacing human judgment with reality-based settlement does not eliminate all risks. A system may optimize predicted reality by changing the environment, manipulating measurements, delaying settlement, exploiting execution infrastructure, or staking only easy-to-settle claims. The authors characterize preventing these behaviors while preserving reliable settlement as an unresolved, safety-critical design problem.

References

This is an open, safety-critical design problem (see Ethics Statement).

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward  (2609.09776 - M et al., 9 Sep 2026) in Section 10, “Failure Modes”; see also Ethics Statement