Scaling of optimal RANDAO manipulation with epoch length ℓ
Determine whether the percentage improvement over the honest policy achieved by the optimal RANDAO manipulation strategy in the reduced Markov decision process M'_G scales with the epoch length ℓ in the manner indicated by the truncated evaluation plot for ℓ ∈ {16, 32, 64, 128}. Specifically, ascertain if the empirical scaling behavior shown in the figure “Percentage improvement over honest for ℓ ∈ {16,32,64,128}” accurately reflects the true dependence on ℓ despite floating‑point numerical instability in the evaluation for ℓ > 32.
References
We conjecture that this plot is representative of how the results scale with ℓ, although unlike our main results the experiments are not provably accurate due to the aforementioned numerical instability. If one desires provable numerical guarantees on these quantities, one would need an analysis of numerical error induced by floating point representations of the machines that run the evaluation.
Three directions remain open. First, \cref{sec:mitigation} analyzes a single tail-slashing curve built from a decaying penalty and a multi-miss multiplier. A natural extension is to explore alternative slashing curves, which may achieve smaller costs for honest validators under the same deterrence guarantee. Second, our framework works at the epoch level and we assume that the adversary commits to an action for the entire epoch at once. However, in practice, the adversary may do even better by adapting their strategy based on the blocks they see in the current epoch online. Analyzing this adaptive strategy is an interesting direction for future work. Third, the empirical gap between our upper and lower bounds is much smaller than our sample-complexity bound predicts, suggesting room for a tighter analysis.