EESE-Pool: Science Benchmark & Pool Testing
- EESE-Pool is a dual-purpose term referring either to a non-public repository of over 100,000 science question-answer instances for benchmarking or to a pool-testing method for epidemiological prevalence estimation.
- The science evaluation repository is built through a rigorous three-stage data engine and refined with expert and LLM-assisted quality control to ensure diverse, representative, and difficulty-calibrated item sampling.
- The epidemiological method uses geometrically spaced pool sizes and maximum-likelihood estimation to accurately determine infection probabilities across a broad dynamic range.
EESE-Pool is a term used for two distinct research constructs. In the context of the Ever-Evolving Science Exam, it denotes a non-public “question–answer” instance repository containing over 100,000 science items and serving as the authoritative master set from which periodically refreshed 500-instance EESE releases are drawn (Wang et al., 22 Jul 2025). In a separate epidemiological context, EESE-Pool denotes “Estimation of Epidemiological Spread via Efficient Pool-testing,” a methodology drawn from variable pool testing for prevalence estimation that mixes pool sizes and reconstructs infection probability by likelihood-based estimation (Bergel, 2020). The shared label therefore refers either to a dynamic science benchmarking substrate or to a pooled-testing estimation procedure, depending on disciplinary context.
1. Terminological scope
The two uses of EESE-Pool differ in domain, object type, and operational purpose.
| Usage | Domain | Core object |
|---|---|---|
| EESE-Pool in EESE | foundation-model evaluation | non-public repository with over 100,000 science question–answer instances |
| EESE-Pool in variable pool testing | epidemiological estimation | methodology using mixed pool sizes and likelihood-based estimation of infection prevalence |
In the first usage, EESE-Pool is the large, non-public reservoir that underwrites leakage-resilient, periodically updated science evaluation slices. In the second, EESE-Pool is a procedure for estimating prevalence over a wide dynamic range without an initial guess of , using pools of geometrically varying sizes and a maximum-likelihood estimator. The overlap is nominal rather than methodological. This suggests that citation by arXiv identifier is necessary when the term appears without expansion.
2. EESE-Pool as the master repository for the Ever-Evolving Science Exam
Within the two-level EESE framework, EESE-Pool is the non-public master set, while EESE is a periodically updated 500-instance subset sampled and validated for low-overhead evaluation (Wang et al., 22 Jul 2025). EESE-Pool contains over 100,000 science instances across 5 disciplines and 500+ subfields. The five major areas are Natural Sciences, Agricultural Sciences, Medical Sciences, Engineering & Technological Sciences, and Humanities & Social Sciences. Natural Sciences is described as including Mathematics, Physics, Chemistry, Biology, Earth & Space Sciences, and related fields.
The repository is designed around the triad of Range, Reach, and Rigor. Range refers to scale: more than 100,000 instances support statistically robust evaluation and long-term stability, and underwrite the ability to draw many distinct 500-item EESE releases while maintaining representativeness. Reach refers to coverage across five broad disciplinary areas, more than 500 granular subfields, and five question formats: single-choice, multiple-choice, fill-in-the-blank, true/false, and open-ended. Rigor refers to methodology: LLM-assisted coarse-grained quality control, expert correction, fine-grained refinement through three parallel branches, and difficulty calibration using aggregated model accuracy with expert overrides.
The role of EESE-Pool is therefore not only archival. It is the mechanism that makes periodic resampling possible while preserving disciplinary breadth, format diversity, and difficulty stratification. No explicit global scoring or coverage formula is given, but difficulty is operationalized by aggregated model accuracy thresholds; for example, if average LLM accuracy on an item is below a threshold , the item is categorized as “hard.” A plausible implication is that the pool functions as a controlled latent distribution from which benchmark instances are repeatedly instantiated.
3. Construction pipeline and refinement framework
EESE-Pool is built by a three-stage Data Engine followed by a fine-grained Data Refinement process (Wang et al., 22 Jul 2025). In the Transcription Stage, over 300 domain experts collect questions and answers from textbooks, public databases, and online sources. All items are standardized into a unified format consisting of metadata, stem, options/answer, and rationale. Coarse-grained quality control then proceeds in two steps: LLMs flag formatting, factual, or logical errors, and experts manually correct flagged items. Each instance is taxonomically labeled as one of 163 subfields under GB/T 13745-2009 and assigned one of the five question formats.
In the Expansion Stage, specialists author new, high-value instances for under-represented subfields, extending coverage to over 500 subfields. All newly authored items undergo the same coarse-grained quality control used in transcription, with the stated goal of maintaining consistency and factual accuracy. In the Categorization Stage, multiple top-tier LLMs answer every instance in zero-shot. Based on aggregated model accuracy, items are assigned to easy, medium, or hard tiers according to predefined score thresholds. Ambiguous or outlier cases are manually reviewed and relabeled by experts.
The subsequent refinement layer is a parallel three-branch framework intended to raise overall difficulty and eliminate trivial items. Enhancement by Distraction uses low human involvement: plausible but incorrect distractors are injected into multiple-choice items, or extraneous details are added to open-ended items; candidates are auto-generated via LLMs and then expert-verified. Enrichment by Cross-Disciplinary Integration uses medium human involvement: contexts or concepts from other fields are introduced so that items require multi-field knowledge synthesis; LLMs provide initial interdisciplinary prompts, which experts refine for precision. Expert-Driven Refinement uses high human involvement: experts manually rewrite or restructure items to embed subtle logical complexity or multi-step reasoning, followed by fine-grained expert validation to preserve scientific rigor.
The dataset statistics reported for the resulting pool include total instances greater than 100,000, five major disciplinary areas, more than 500 subfields, five question formats, and a difficulty distribution stratified into easy, medium, and hard tiers, with refinement shifting mass toward medium and hard. This architecture places EESE-Pool closer to a continuously maintained evaluation substrate than to a static benchmark release.
4. Sampling, representativeness, and empirical behavior
Each EESE release consists of 500 instances drawn at random from EESE-Pool subject to discipline, subfield, format, and difficulty quotas inherited from the pool (Wang et al., 22 Jul 2025). The principal rationale is leakage resilience: by regularly rotating the 500 items, static memorization by models is minimized. A second rationale is efficiency: a 500-item test set reduces compute and cost overhead while maintaining strong correlation with the full pool.
Representativeness is evaluated using Spearman rank-order correlation coefficients (SRCC) between model rankings on EESE and on EESE-Pool. The reported SRCC heatmaps have diagonals close to 1, which is presented as confirmation that the sampled EESE slices preserve ranking fidelity relative to the full repository. This makes EESE a proxy benchmark derived from a much larger benchmark substrate rather than an independently curated small test set.
The empirical findings reported for the pool emphasize heterogeneity across disciplines and item difficulty. Evaluation on the 100K+ pool reveals large performance swings across fields; models may excel in some STEM subfields while lagging in humanities. All representative LLMs show lower accuracy after refinement on hard items, which is taken as confirmation that the three-branch framework successfully raises difficulty. Experiments on 32 open- and closed-source models indicate that EESE differentiates strengths and weaknesses in both scientific fields and cognitive dimensions. At the same time, state-of-the-art proprietary LLMs are reported to trail human experts by a wide margin, underscoring the pool’s intended rigor in challenging scientific reasoning.
5. EESE-Pool as variable pool testing for infection spread estimation
In the epidemiological usage, EESE-Pool is a three-stage procedure consisting of pool-size design, testing, and estimation (Bergel, 2020). A total budget of tests and a target dynamic range are first chosen. The smallest pool size is set so that , and a logarithmic spacing factor is used so that pool sizes grow geometrically and satisfy . The stated purpose is to guarantee pools small enough to detect high and large enough to detect very low . In the testing stage, the 0-th pool contains 1 distinct specimens, each pool is assayed by a standard binary test such as RT-qPCR, and the outcome is encoded as
2
The probabilistic model assumes that each sample is independently infected with probability 3. For a pool of size 4,
5
Equivalently,
6
Given outcomes 7 and pool sizes 8, the joint likelihood is
9
The estimate 0 is obtained by numerical root-finding, for example binary search, of a monotonic equation. In the homogeneous special case 1,
2
A central design claim is that geometric spacing of pool sizes over 3 ensures that for any true 4, some pools will have size roughly 5. The method therefore operates without an initial guess of 6. The documented spacing tradeoff is that larger 7 covers a wider range in fewer tests but sacrifices local accuracy, whereas smaller 8 improves local resolution at the cost of range. Typical choices given are 9, 0–1, and 2–3, yielding coverage from 4 to 5.
Performance is summarized using the relative root-mean-square error
6
With 7, 8, and 9, so that 0, the method attains 1 for all 2 using approximately 3 total samples. Restricting total samples to 1,000 and adjusting to 4 retains 5 down to 6. With 7 tests and 8, the method covers 9 with a 95% confidence band of roughly 0; with 1 and 1,000 samples, it remains informative above 2.
The assumptions and limitations are explicit. The model assumes a perfect binary test with no false positives or false negatives, an independent infection model in which samples are i.i.d. with infection probability 3, sufficient sample volume even for very large pools, and manageable logistical overhead for preparing many pools of varying sizes. The text also notes that pooling dilution can increase error rates and that future work should incorporate sensitivity and specificity curves as a function of pool size 4.
6. Distinctions, interpretive boundaries, and citation practice
The benchmark usage and the epidemiological usage share a label but differ at every substantive level. One is a non-public science question–answer repository with more than 100,000 instances, a three-stage data engine, fine-grained refinement, and periodic 500-instance benchmark sampling (Wang et al., 22 Jul 2025). The other is a prevalence-estimation method based on geometrically spaced pool sizes, a binomial observation model, and maximum-likelihood reconstruction of 5 from binary pooled tests (Bergel, 2020).
A plausible source of confusion is that both usages emphasize efficiency and robustness, but the corresponding meanings are domain-specific. In the science-benchmark setting, efficiency refers to low-overhead evaluation through a 500-item sample that preserves ranking fidelity to the full pool. In the epidemiological setting, efficiency refers to estimating prevalence accurately with a small number of tests and mixed pool sizes over a broad dynamic range. The benchmark version is concerned with leakage resilience, breadth of coverage, and expert curation; the epidemiological version is concerned with dynamic-range coverage, likelihood estimation, and test-budget allocation.
The distinction is therefore not merely terminological. EESE-Pool in the evaluation literature is an infrastructural repository for continuous benchmarking of foundation models’ scientific reasoning. EESE-Pool in pooled-testing methodology is an estimation framework for infection spread. This suggests that the unexpanded acronym is insufficiently specific in cross-disciplinary settings, and that the associated paper title or arXiv identifier is the most reliable disambiguator.