---
title: 'Hedge-Bench 1.0: AI Financial Analysis Benchmark'
url: https://www.emergentmind.com/topics/hedge-bench-1-0
type: topic
---

# Hedge-Bench 1.0: AI Financial Analysis Benchmark

Hedge-Bench 1.0 is an open-source benchmark specifically designed for evaluating AI agents on hard, realistic financial reasoning tasks that mirror professional hedge fund analyst workflows. Unlike question–answer or span-matching datasets prevalent in financial NLP evaluation, Hedge-Bench 1.0 centers on multi-step, evidence-grounded, and synthesis-focused tasks, assessing both outcome correctness and the explicit reasoning process by which investment judgments are produced [2606.03918].

## 1. Benchmark Definition and Motivation

Hedge-Bench 1.0 comprises 102 environments, each modeled after real on-the-job tasks encountered by hedge fund analysts. Each environment presents agents with a sandbox of relevant primary documents—such as SEC filings, earnings call transcripts, peer company financials, and curated news articles—frozen to the task date. The agent is given an open-ended instruction (e.g., "Assess whether Iridium’s L-band spectrum confers a durable moat in government markets over the next 5 years; take a position, cite sources.") and must decompose this into sub-questions, perform relevant analyses, and synthesize a final investment view, with explicit source citations throughout.

The benchmark was created to address significant limitations in prior financial QA evaluations, which mostly consist of fact retrieval, numerical calculations, or span selection, and thus fail to capture the open-ended, multi-step argumentation found in authentic analyst workflows. Prior resources such as FinQA, TAT-QA, DocFinQA, and FinanceBench focus on low- to mid-level skills (e.g., extracting numbers from tables, running formulas), but Hedge-Bench uniquely measures high-leverage judgment, sub-task decomposition, and evidence synthesis [2606.03918].

## 2. Dataset Construction and Expert Reasoning Traces

The tasks in Hedge-Bench 1.0 were sourced over nine months from real hedge-fund analyst workflows. Out of 5,112 candidate tasks (and over 20,000 sub-tasks), two independent reviewers selected the 102 most challenging and highest-fidelity for final inclusion. For each, a closed set of source documents is bundled per environment and exposed via a Dockerized container with a CLI harness (Harbor).

Crucially, for grading and reference, reasoning traces were collected by pairing two professional analysts in a real-time, phone-based walk-through of each task. These discussions were transcribed and aligned to the supporting document pool, from which a rubric was extracted:
- Each environment is structured by T themes (lines of inquiry).
- Each theme is partitioned into n required moves (sub-themes), with explicit references to supporting documents for each.
This one-trace-per-task rubric serves as the sole ground truth for coverage and reasoning validation, capturing both the logic and the document grounding of expert financial analysis.

## 3. Evaluation Methodology

Hedge-Bench 1.0 uses a deterministic, rubric-driven grading pipeline that consists of several LLM-constrained sub-checks:

1. **Grounding Check:** All factual claims (numbers, dates, quotes, entities) in the agent’s output must be exactly verified against cited source files. Any claim without exact evidence is flagged, and dependent reasoning moves are marked “tainted.”
2. **Coverage Check:** The agent’s answer is mapped onto the rubric structure, and for each theme with n moves, it is considered "covered" once τ = max(1, min(n–1, 3)) required moves are valid and untainted.
3. **Synthesis Check:** To achieve the highest score, the answer must contain at least one explicit synthesis—meaning the reconciliation of opposing data points into a unified conclusion.

Rubric-based scoring assigns a dense score, $s \in \{0,1,2,3,4\}$, to each trial:
- 4: all themes covered and synthesis present
- 3: all themes covered, no synthesis
- 2: at least two themes covered
- 1: at least one theme covered
- 0: less than one theme covered

Two main evaluation metrics are reported:
- **Mean Dense Score:** The mean assigned score across all valid trials and environments.
- **Sparse Pass@1:** The proportion of trials (per environment, then macro-averaged) attaining the top rubric score ($s=4$). Confidence intervals are estimated as $1.96 \cdot \sigma/\sqrt{N}$, $N$ being the number of environments.

## 4. Task Taxonomy and Reasoning Types

Hedge-Bench’s suite of tasks spans six principal categories:
1. Valuation
2. Growth & Expansion
3. Mergers & Acquisitions (M&A)
4. Competitive Positioning
5. Operational Execution & Strategy
6. Risk

Task styles include document ingestion & retrieval, fact-checking & source-grounding with inline citations, numerical computation (e.g., multiples, earnings normalization), sub-task decomposition, and explicit open-ended synthesis (e.g., scenario analysis, reconciling contradictory evidence).

Each task is constructed to assess agent ability to map realistic, multi-theme instructions into structured investigations reflecting senior analyst practice. For example, a valuation task may require parsing multiple filings, running comparative metrics, analyzing management commentary, and explicitly weighing risks.

## 5. Baseline and Frontier Model Performance

The following table summarizes model performance as detailed in [2606.03918]:

| Model               | pass@1 (%) | mean_dense |
|---------------------|:----------:|:----------:|
| Claude-Sonnet-4.6   |   15.5     |   1.92     |
| Claude-Opus-4.7     |   14.2     |   1.84     |
| GPT-5.5             |   12.7     |   1.68     |
| Gemini-3.5-Flash    |   11.8     |   1.68     |
| Claude-Opus-4.8     |   10.6     |   1.62     |
| Claude-Haiku-4.5    |    4.1     |   1.24     |
| Gemini-3.1-Pro      |    3.2     |   1.12     |
| GPT-5.4-Mini        |    0.8     |   0.75     |

Key observations include:
- The best available models resolve fewer than 1 in 6 tasks end-to-end and exhibit substantial difficulty in judgment-heavy domains (e.g., Risk, M&A, Competitive Positioning), where pass@1 is often <2%.
- Hallucination rates are high, with 80–90% of model trials containing at least one ungrounded claim, posing significant issues for reliable deployment.
- Common failure modes include incomplete coverage of rubric moves (especially when τ>1), omissions in synthesizing opposing data, grounding errors such as fabricated numbers or misattributed quotes, and shallow decomposition (completing only 1–2 reasoning steps).

## 6. Benchmark Distribution, Workflow, and Integration

Hedge-Bench 1.0 is released under an open-source license at https://github.com/trata/hedge-bench. For each environment:

- Directory structure includes source data files, the extracted rubric, Dockerfile for environment isolation, and a prompt instruction.
- Evaluation relies on the Harbor CLI harness (`pip install harbor-bench`) supporting both Python and command-line interfaces.
- Users can instantiate the harness, load environments, and run agent trials with result parsing for pass@1 and mean_dense metrics.

Example workflow (Python API):

```python
from harbor import HarborHarness, TaskEnv
from my_llm_agent import MyAgentClient

harness = HarborHarness(model="gpt-5.5", trials=8, timeout=300)
agent   = MyAgentClient()

for env_path in sorted(Path("hedge-bench-1.0").glob("env_*")):
    env  = TaskEnv(env_path)
    res  = harness.run(env, agent)
    print(f"{env.name}: pass@1={res.pass1:.1%}, dense={res.mean_dense:.2f}")
```

This deterministic, rubric-grounded grading procedure provides reproducibility across researchers and institutions.

## 7. Limitations and Future Directions

Hedge-Bench 1.0’s rubric structure derives from a single analyst-pair transcript per task; it is possible that alternative deconstructions might exist, and the one-shot rubric extraction can deviate from rich expert discussions. Robustness currently depends on correct LLM judging, making the pipeline temporarily vulnerable to judge errors (trials with parsing errors are zeroed out). Furthermore, context window limitations bound the number and length of input documents per environment.

Planned enhancements for v2 include consensus-based rubrics (multiple analyst perspectives), per-move human validation, systematic re-review of multi-model disagreements, and task expansion to keep pace with rapidly advancing agent capabilities.

The benchmark enables detailed diagnostics of model weaknesses across grounding, synthesis, and coverage, facilitates training using expert reasoning traces as supervision (e.g., in fine-tuning or RLHF), and supports practical vetting for real-world financial analysis, due diligence, and risk assessment [2606.03918].

Source: https://www.emergentmind.com/topics/hedge-bench-1-0