---
title: Arena-Hard-200 Benchmark
url: https://www.emergentmind.com/topics/arena-hard-200-benchmark
type: topic
---

# Arena-Hard-200 Benchmark

Arena-Hard-200 is a curated, highly challenging benchmark designed to assess the advanced reasoning and generalization abilities of large language models (LLMs). Constructed via the ArenaBencher automatic evolution framework, Arena-Hard-200 extends the GSM8K mathematical problem-solving test set by targeting emerging LLM failure modes, increased difficulty, and robust resistance to data contamination. As a distillation of the hardest items synthesized through multi-model adversarial evaluation, iterative LLM augmentations, and rigorous validity checks, Arena-Hard-200 serves as a high-fidelity stress test for differentiating next-generation LLMs on mathematical reasoning and related domains [2510.08569].

## 1. Framework and Method

Source: https://www.emergentmind.com/topics/arena-hard-200-benchmark