AGGBench: Aggregation Query Benchmark
- AGGBench is a benchmark that assesses systems performing entity-level aggregation queries, focusing on exhaustive evidence retrieval and strict completeness.
- It combines a core corpus of research papers with noise documents to simulate realistic, large-scale challenges for multi-chunk evidence identification.
- The evaluation employs metrics like chunk-level coverage, ACE, and NACE to provide actionable insights into system performance in terms of recall and accuracy.
AGGBench is a benchmark designed to evaluate the completeness and evidence coverage of systems performing entity-level aggregation queries over unstructured text corpora. Unlike typical question answering tasks, aggregation queries require systems to exhaustively identify all entities satisfying complex, compositional conditions, thereby placing stringent demands on evidence retrieval, disambiguation, and aggregation processes. AGGBench is corpus-bounded, explicitly prohibiting the use of external knowledge and focusing evaluation on the ability to “find all” such entities within realistic, large-scale, noisy corpora (Zhu et al., 1 Feb 2026).
1. Formalization of Aggregation over Unstructured Text
AGGBench targets the formal problem of entity-level aggregation querying under a strict completeness regime. The corpus is partitioned into text chunks; denotes all entity mentions across . Each query specifies:
- an entity type
- a predicate set , each a boolean condition on entities of type , satisfied only if there is explicit evidence in
The exact answer set is:
where 0. For each 1, at least one supporting chunk evidencing each predicate must be found. AGGBench emphasizes strict recall: all satisfying entities and corresponding evidentiary chunks must be recovered, not merely plausible answers.
The key evaluation metric is evidence coverage at the chunk level:
2
where 3 is the gold set of evidence chunks, and 4 is the set returned by the system.
2. Benchmark Construction and Annotation
The construction of AGGBench is designed to enable completeness-oriented evaluation under realistic, noisy conditions.
Corpus Design
- Core corpus: 45 research papers from the “graph retrieval–augmented generation” literature (e.g., NeurIPS, ICLR), chunked into 200–300-token segments, totaling 4,755 chunks.
- Expansion: 11,539 unrelated, noise-inducing documents were added, yielding a final corpus of 16,294 chunks. BM25 proximity filtering ensured that no new satisfying entities were introduced by noise documents.
Query and Condition Generation
- Entity types (5) were extracted and ranked by frequency, with manual curation to remove ambiguous categories.
- Conditions (6) were mined via high-frequency descriptive phrases (e.g., “used for multi-hop QA,” “applied to legal domain”), then manually refined for compositionality and unambiguous meaning.
- Resulting queries are natural-language prompts about entity counts, such as:
- “How many datasets are used for multi-hop question answering?”
- “How many papers apply to the legal domain?”
- The benchmark comprises 362 queries: 100 base (single-condition), and 262 composite (multiple AND/OR conditions).
Evidence Annotation Workflow
Annotation is two-stage:
- LLM pre-annotation: A LLM annotates each (query, chunk) pair as positive/negative and extracts candidate entities, filtering 90% as clear negatives.
- Human verification: Annotators review and correct LLM outputs, ensure adequate evidence grounding for each entity, and consolidate multi-chunk evidence. Only about 10% of LLM annotations require correction.
3. Metrics and Evaluation Protocol
AGGBench provides a modular evaluation protocol targeting both completeness and accuracy:
- Evidence completeness is measured by chunk-level recall:
7
- Result accuracy metrics:
- ACE (Absolute Count Error): 8, where 9 (gold count), 0 is the system’s count.
- NACE (Normalized ACE): 1, with 2 preventing division by zero.
Coverage captures whether all relevant evidence is found. High ACE/NACE generally reflects low coverage, underscoring the principal challenge of achieving exhaustive retrieval rather than just plausible responses.
4. Dataset and Implementation Resources
AGGBench is distributed with both raw and processed data, as well as modular code for evaluation and agentic baseline experiments:
- Repository structure:
data/raw_core/: original PDFs/texts of core papersdata/chunks/: tokenized chunk filesdata/queries.json: full query set with predicate templatesdata/gold_answers.json: gold-standard entity lists and chunk evidence mappingscode/benchmark.py: harness for evaluation and scoringcode/chunk_retriever.py: BM25 and dense retriever implementationscode/dfa_agent.py: DFA agent baselinecode/utils/: tokenization and chunking scriptsrequirements.txt: dependencies, including transformers, faiss, and rank_bm25
- Installation and usage: Python 3.9+ required, with setup via
pip install -r requirements.txt. Data and code are downloaded and referenced by setting the\mathcal{E}(C)$9 - Access: Data and code are available at https://anonymous.4open.science/r/DFA-A4C1 (Zhu et al., 1 Feb 2026).
5. Benchmark Statistics
AGGBench is characterized by its scale, evidence density, and compositional query types.
| Statistic | Value/Range | Notes |
|---|---|---|
| Total queries | 362 | 100 base (single-condition), 262 composite |
| Answer set size ($\mathcal{E}(C)$3) | 165 queries with $\mathcal{E}(C)$4 answers | Max: 20 (single), 29 (composite) |
| Core corpus | 45 docs → 4,755 chunks | 294 gold-evidence chunks (6.18% evidence density) |
| Expanded corpus | 16,294 chunks | 178 gold-evidence chunks (1.09% density) |
| Evidence per query (avg.) | $\mathcal{E}(C)$5 chunks | Varies by query; reflects multi-chunk evidence necessity |
| Query compositionality | 228 double, 34 triple conditions | 42 AND, 220 OR queries |
This evidentiary sparseness, with many queries requiring synthesis of 8 or more distinct chunks, reflects the realistic difficulty of the “find-all” aggregation setting in unstructured text.
6. Comparison to Prior Approaches
AGGBench exposes shortcomings in prevalent methods for QA over text:
- Text-to-SQL (schema-first approaches):
- Depend on brittle extraction pipelines to convert text into structured DBs, often yielding limited coverage.
- Fixed schemas prevent on-the-fly synthesis of new, compositional natural language queries.
- Even accurate SQL cannot query for entities missed by initial extraction, undermining completeness.
- Retrieval-Augmented Generation (RAG, rank-then-read):
- Scoring functions optimize top-$\mathcal{E}(C)$6 relevance, not exhaustive recall. When $\mathcal{E}(C)$7, many evidence chunks are omitted and answer counts are under-reported.
- Increasing $\mathcal{E}(C)$8 undermines answer precision by flooding context with irrelevant or noisy chunks.
AGGBench is explicitly designed to isolate aggregation-specific failure modes, such as ambiguous entity boundaries, errors in predicate application (“filter roll-backs”), and the proper alignment of multi-chunk evidence. Evaluation protocol is centered on recall and completeness, rather than mere plausibility or relevance (Zhu et al., 1 Feb 2026).
7. Applications and Limitations
Use Cases
- Legal e-discovery and contract analytics: e.g., “Find all contracts/papers that mention clause X.”
- Financial and compliance auditing: e.g., “How many companies exhibit risk-factor Y in their disclosures?”
- Investigative journalism: e.g., “List all sources meeting conditions A∧B among thousands of documents.”
- Data-analysis agents: for scenarios where exhaustive filtering of entities from large text corpora is required.
Limitations
- Domain specificity: The core is limited to research papers in the graph RAG field; adaptation to domains such as law or finance mandates new data curation and annotation.
- Query scope: Only entity-count (aggregation) queries are supported; AGGBench does not address sum, average, or other numerical aggregations beyond counting.
- Ambiguity handling: C-type entity ambiguities (granularity, deduplication, unknown labels) are rare in AGGBench and only addressed qualitatively.
- No external knowledge: The protocol enforces a strict corpus-only (no outside KBs) policy.
- Annotation cost: Despite initial LLM filtering, manual correction remains necessary for approximately 10% of labels.
AGGBench thus provides a rigorously defined foundation for testing, diagnosing, and benchmarking completeness-oriented aggregation query methods over unstructured text, with a reference agentic baseline (the DFA agent) that modularizes the disambiguation, filtering, and aggregation pipeline and exposes key system-level failure points (Zhu et al., 1 Feb 2026).