Replication across multiple Arena worlds

Aggregate FM-Bench Arena results across several shared worlds using a rank-based aggregate of per-world finishes, thereby determining whether the reported Arena ordering is robust beyond the single seed-7 world.

Background

The Arena results are based on one shared seed-7 world with no error bars. The paper explicitly cautions that rankings separated by only a few points should be interpreted as ties and notes that a broader evaluation would require aggregating results across multiple Arena worlds.

The unresolved methodological problem is therefore to run and combine several shared-world competitions in a principled way, using ranks within each world rather than directly pooling scores that may depend on world-specific conditions.

References

The Arena (\S\ref{sec:arena}) is one shared seed-7 world with no error bars, so its orderings within a few points should be read as ties, and aggregating several Arena worlds would call for a rank-based aggregate over per-world finishes, which we leave to future work.

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents  (2608.18423 - Wang et al., 19 Aug 2026) in Appendix, Section “Limitations,” first bullet