---
title: 'RAGEval: RAG Evaluation Framework'
url: https://www.emergentmind.com/topics/rageval-framework
type: topic
---

# RAGEval: RAG Evaluation Framework

Retrieval-Augmented Generation (RAG) evaluation frameworks address the critical need for robust, scenario- and domain-specific benchmarking of RAG systems in settings where real-world documents and specialized knowledge bases are involved. RAGEval is a framework designed for systematic, schema-driven evaluation and benchmarking of RAG pipelines, enabling the rigorous assessment of factual completeness, hallucination rates, and irrelevance in model outputs, and providing high-fidelity synthetic data for both retrieval and generation tasks [2408.01262].

## 1. Schema-Based Pipeline Architecture

RAGEval operates through a four-stage schema-based pipeline:

1. **Schema Summarization**: Beginning with a pool of seed, domain-specific documents (e.g., financial, medical, or legal records), the framework interacts with an LLM to extract the essential "slots" representing entities, values, and events required for domain coverage. Iterative prompting and human-in-the-loop refinement yield a formalized schema $\mathcal S$ representing the data structure.
   
2. **Configuration Generation & Document Synthesis**: Schema $\mathcal S$ is instantiated across diverse configurations $\{\mathcal C_i\}$ using a hybrid rule-based (for structured fields) and LLM-based (for narrative fields) approach. Each configuration becomes a synthetic document $\mathcal D_i$ generated in target style.

3. **QRA Generation (Question–Reference–Answer Triples)**: For each synthetic document, the system prompts an LLM to create diverse question–initial answer pairs $(q_j, a^{(0)}_j)$. It then automatically extracts minimal supporting reference spans $\{r_{j,k}\}$ and enforces answer grounding by aligning each fact in $a^{(0)}_j$ to evidence in $\mathcal D_i$. Final keypoints $K_j$ are distilled from each refined answer $a_j$.

4. **Human and Automated Quality Verification**: Human annotators validate a random subset of data for factual correctness and document quality. Automated metrics are cross-validated against these human judgments to ensure reliability and stability.

This schema-to-data process enables RAGEval to efficiently build domain-meaningful, large-scale evaluation datasets suitable for real-world RAG settings, as summarized by the following pipeline:

```python
def RAGEval_pipeline(seed_docs):
    S = extract_schema_with_LLM(seed_docs)
    configs = generate_configs(S)
    documents, qra_triples = [], []
    for C in configs:
        D = generate_document_from_config(C)
        Q, A0 = LLM_generate_QA(D, C)
        R = extract_references(D, Q, A0)
        A = refine_answer_with_refs(A0, R, D)
        K = extract_keypoints(A)
        qra_triples.append((Q, R, A, K))
    return documents, qra_triples
```
[2408.01262]

## 2. Formal Definitions of Core Metrics

RAGEval introduces three principal metrics for evaluation at the fine-grained, keypoint level. Given a QA pair $(q,a)$ and its reference keypoints $K = \{k_1, \dots, k_n\}$, these are:

- **Completeness**:
  $$
  \mathrm{Comp}(a, K) = \frac{1}{|K|} \sum_{i=1}^n \mathbbm{1}[a \text{ semantically covers } k_i]
  $$
  Fraction of required facts present in the answer.

- **Hallucination**:
  $$
  \mathrm{Hallu}(a, K) = \frac{1}{|K|} \sum_{i=1}^n \mathbbm{1}[a \text{ contradicts } k_i]
  $$
  Fraction of facts in $K$ that are contradicted (i.e., clear hallucinations).

- **Irrelevance**:
  $$
  \mathrm{Irr}(a, K) = 1 - \mathrm{Comp}(a,K) - \mathrm{Hallu}(a,K)
  $$
  Fraction of required facts not present and not contradicted.

Retrieval metrics include:

- **Recall**:
  $$
  \mathrm{Recall} = \frac{1}{N} \sum_{i=1}^N \mathbbm{1}[G_i \subseteq \cup_j R_j]
  $$
  Proportion of instances where all ground-truth references $G_i$ are contained in retrieved passages $R_j$.

- **Effective Information Rate (EIR)**:
  $$
  \mathrm{EIR} = \frac{\sum_{i:\, G_i \text{ matched}} |G_i \cap R_\mathrm{all}|}{\sum_j |R_j|}
  $$
  Fraction of retrieved tokens that are genuinely informative for the reference.
[2408.01262]

## 3. Dataset Construction and Validation

The framework, as instantiated in the DragonBall dataset, covers finance, law, and medical domains across Chinese and English languages. For each domain:

- Documents per language: 40 finance, 30 legal, 38 medical.
- Generated QRA triples: 6,711, encompassing factual, multi-hop, summarization, info-integration, numeric comparison, time-series, and unanswerable types.
- Benchmarks against zero-shot and one-shot document generation approaches (single-prompt or with one example), with RAGEval outperforming both in human ratings for safety, clarity, conformity, and richness.

Comprehensive human annotation was performed on 420 random QRA triples across domains/languages. RAGEval’s automated metric scores for Completeness, Hallucination, and Irrelevance showed absolute deviations of less than 0.015 from human judgment, confirming high consistency.
[2408.01262]

## 4. Experimental Results and Retrieval/Generation Benchmarks

Retrieval experiments employed BM25, GTE-Large, BGE-Large, and BGE-M3 (for both Chinese and English), while generation tested MiniCPM-2B-sft, Baichuan-2-7B-chat, Qwen1.5/2, Llama3-8B-instr, GPT-3.5-Turbo, and GPT-4o.

Key quantitative findings include:

- **Generation Results**: GPT-4o achieved the best Completeness ($0.52$ CN, $0.68$ EN) and lowest Hallucination, but competitive performance was observed in Qwen1.5-14B and Llama3-8B.
- **Retrieval**: BGE-M3 excelled in Chinese (Recall $= 0.84$, EIR $= 0.05$) with the best downstream Completeness and Hallucination. GTE-Large led in English.
- **Hyperparameters**: TopK sweeps (2–6) increase recall and completeness, with chunking strategies yielding domain- and language-specific best practices (smaller, more chunks for Chinese; moderate-size, fewer for English).
[2408.01262]

## 5. Algorithmic and Implementation Principles

The hybrid schema-to-data paradigm unifies rule-based strictness for structured fields with LLM-driven diversity for narrative fields. Each configuration $\mathcal C_i$ is deterministically mapped to a synthetic document and a set of QRA triples (including prominence of answer grounding and minimal hallucination risk).

Released code supports:

- Data and QRA artifact management (`data/dragonball/`, `configs/`)
- Modular shell scripts for each pipeline stage (extraction, config generation, document synthesis, QRA generation, retrieval/generation evaluation)
- Implementation of metrics and unit tests (`src/metrics.py`)
- Human annotation templates and guidance (`human_eval/`)
[2408.01262]

## 6. Limitations and Prospects for Extension

RAGEval is reliant on the foundational quality of the prompting LLM for both schema extraction and keypoint mining, necessitating manual schema cleaning and answer correction if the base model exhibits high hallucination rates. The dataset construction methodology is currently validated on finance, law, and medical domains; generalization to domains with high narratological or non-tabular complexity may require further adaptation (such as unsupervised schema induction).

The granularity of evaluation (via Comp/Hallu/Irr) depends on accurate keypoint extraction, suggesting that future enhancements using improved semantic overlap measurement may refine these metrics. Further, the semi-automation of schema and answer refinement is labor-intensive, motivating exploration into active learning or more refined LLM-driven filtering.

The general framework and its metrics provide a blueprint for scalable, high-fidelity RAG evaluation in both academic and industrial environments, and open avenues for further benchmarking and pipeline tuning across verticals [2408.01262].

---

**References**:  
- "RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework" [2408.01262]

Source: https://www.emergentmind.com/topics/rageval-framework