Benchmarking BFT protocols under actual Byzantine behavior

Develop and validate experimental methods for benchmarking and testing Byzantine fault-tolerant protocols under actual Byzantine behavior.

Background

The paper evaluates the adaptive leader-count scheduler using simulated network degradations rather than actual Byzantine behavior. It explicitly identifies experimental benchmarking under adversarial Byzantine actions as an unresolved methodological problem. Formal proofs can establish worst-case safety and liveness guarantees, but they do not replace empirical evaluation of protocol behavior under realistic Byzantine strategies.

References

Benchmarking and testing BFT protocols under actual Byzantine behavior is an open problem; the state of the art establishes worst-case guarantees through formal proofs, which we give in \Cref{sec:proofs} for both safety and liveness.

Barnacle: Adaptive Multi-Leader Scheduling for DAG-Based Consensus  (2609.03978 - Angeli et al., 3 Sep 2026) in Section 5, Evaluation