Detection-risk versus information trade-off under threshold geometry

Determine whether the difference in threshold geometry between BB84 and E91 dictates how an adaptive attacker trades detection risk for information gain by sweeping the detection penalty in a fine-grid dynamic programme.

Background

The reinforcement-learning experiments compare strict zero-detection policies with policies trained using a softer detection penalty. On BB84, allowing some detection improves the learned policy's information, but the dynamic-programming upper bound is attainable without detection; on E91, the strict and tolerated agents coincide. The two protocols also differ in how their abort thresholds relate to the channel: BB84's threshold is chosen independently of the channel, whereas E91 approaches a noise level at which the honest CHSH value itself reaches the operational threshold.

The paper leaves unresolved whether this geometric difference is responsible for the distinct observed trade-offs between detection risk and extracted information. A fine-grid dynamic-programming sweep over detection penalties is proposed as the means of resolving the question.

References

It remains untested whether this difference in threshold geometry dictates how an attacker trades detection risk for information gain. Sweeping the detection penalty within a fine-grid dynamic programme would resolve this question.

— Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift  (2610.01792 - Mordarski et al., 1 Oct 2026) in Section 5.6, subsection “Tolerance for detection”