Monotonicity of optimal-action probability under ABRA
Prove that, for any fixed noise parameter β, the probability of selecting an optimal action under the stationary distribution concentrated on an optimal recurrent class of the Approximate Best Response Algorithm increases with the rationality parameter p.
References
Based on this observation, we make the following conjecture whose proof is the left for the future work. The probability of choosing an optimal action q*_p under stationary distribution \pi*_p concentrated on an optimal class increases with rationality parameter p for any fixed noise parameter \beta.
Based on this observation, we make the following conjecture whose proof is the left for the future work. The probability of choosing an optimal action $q*_p$ under stationary distribution $\pi*_p$ concentrated on an optimal class increases with rationality parameter $p$ for any fixed noise parameter $\beta$.