Papers
Topics
Authors
Recent
Search
2000 character limit reached

BoostQA: 100B-Token QA Dataset

Updated 18 July 2026
  • BoostQA is a large-scale synthetic QA dataset designed to overcome scale, diversity, and difficulty limitations in traditional QA corpora.
  • It employs a four-stage pipeline—seed curation, dual annotation, two-way synthesis, and answer refinement—to generate challenging, STEM-focused questions.
  • When blended with KnowEdu in a mid-training regime, BoostQA achieves up to a 12.74% improvement on benchmarks, establishing state-of-the-art performance.

Searching arXiv for the specified paper to ground the article in the current record. BoostQA is a 100 billion-token, large-scale synthetic question-answering dataset introduced in "Large-Scale Diverse Synthesis for Mid-Training" (Zhang et al., 2 Aug 2025). It is designed to address three limitations in existing QA corpora for LLMs: insufficient scale, limited knowledge diversity—especially in cross-discipline settings—and the tendency of synthetic pipelines to generate lower-difficulty items. BoostQA is produced by a four-stage pipeline comprising seed curation, dual annotation, two-way synthesis, and answer refinement, and it is consumed in a mid-training stage situated between conventional pre-training and fine-tuning. In the reported setting, a 1:1 blend of BoostQA and KnowEdu yields an average improvement of 12.74%12.74\% on MMLU and CMMLU for Llama-3 8B and establishes SOTA average performance across 12 benchmarks (Zhang et al., 2 Aug 2025).

1. Problem setting and design objectives

The central motivation for BoostQA is the scarcity of high-quality, knowledge-intensive training data for LLMs. Traditional corpora are described as providing limited information, while previous work on synthesized QA is characterized as improving model performance but facing persistent problems in QA data scalability and knowledge diversity, particularly in cross-domain contexts (Zhang et al., 2 Aug 2025).

Three objectives organize the design. The first is scale: publicly available synthetic QA datasets rarely exceed a few tens of billions of tokens, whereas BoostQA expands the scale to 100B tokens. The second is diversity: traditional pipelines typically “rephrase and extend” narrow-domain passages, which yields limited cross-discipline reuse. The third is difficulty: probe experiments on Llama-3 8B checkpoints from 2T to 10T tokens, using a human-calibrated difficulty scorer, show steep degradation from easy to hard tiers, with accuracy falling from 60.7%60.7\% on H1 to 33.8%33.8\% on H5. This diagnosis motivates an explicit effort to amplify hard examples (Zhang et al., 2 Aug 2025).

In this formulation, BoostQA is not merely a larger synthetic corpus. It is structured as an intervention on identified weaknesses in STEM disciplines and high-difficulty data. This suggests that the dataset is intended to modify the composition of training signals rather than simply increase token count.

2. Corpus composition, coverage, and annotation

BoostQA draws its seeds from heterogeneous sources: public QA datasets such as MathQA, high-quality textbooks, and specialized web crawls such as MegaMath-Web (Zhang et al., 2 Aug 2025). Before synthesis, the pipeline performs decontamination using exact 10-gram matching plus embedding-based similarity filtering against all evaluation benchmarks. This places benchmark isolation at the corpus-construction stage rather than treating it as a downstream evaluation concern.

Two annotation systems drive both synthesis and analysis. A discipline classifier with 62 categories is implemented via Qwen2.5-7B-Instruct and is distilled from DeepSeek-R1. A separate difficulty scorer assigns H1–H5 tiers and is calibrated to human pass rates under one-hour tests by top-100 QS students. These labels determine synthesis direction and are also used for distributional analyses (Zhang et al., 2 Aug 2025).

Attribute Reported value
Total size 100 billion tokens
Discipline coverage 62 primary academic disciplines
Top discipline shares Mathematics 51.89%, Computer Science 10.87%, Clinical Medicine 6.25%
Grade levels in synthetic data High School, College, Graduate
Difficulty distribution H1/H2 = 56.86%, H3 = 15.59%, H4/H5 = 27.55%

The discipline profile is explicitly skewed, with Mathematics comprising 51.89%51.89\% of the corpus, followed by Computer Science at 10.87%10.87\% and Clinical Medicine at 6.25%6.25\%. The synthetic data is organized by grade levels—High School, College, and Graduate—but the reported pipeline also observes a down-shift in actual alignment and compensates by prompting one tier higher. A plausible implication is that nominal prompt role and realized item difficulty are not identical, which is why explicit post hoc difficulty control remains necessary.

3. Four-stage synthesis and refinement pipeline

The pipeline begins with seed curation, followed by discipline and difficulty annotation, then two-way synthesis with DeepSeek-R1, and finally answer refinement with DeepSeek-V3 (Zhang et al., 2 Aug 2025). The four-stage decomposition is central to the identity of BoostQA and distinguishes it from approaches that only extend pre-existing QA pairs.

The first synthesis branch is STEM-focused multi-grade synthesis. It uses role prompts corresponding to “High School,” “College,” and “Graduate” educators. Prompting enforces format constraints; for example, multiple-choice items must contain four distractors, while essay questions require a self-contained prompt plus an explanatory solution. The output size is configurable, typically n=10n=10 per seed, and the seed set is over-sampled toward STEM. This branch yields broad coverage but tends toward lower-difficulty items.

The second branch is high-difficulty synthesis, described as a Difficulty Booster. Here the role is fixed to “Graduate,” and an integrated difficulty validator rejects any item with estimated pass rate ≥30%\ge 30\%. Sampling weights are re-balanced so that seeds labeled H4/H5 receive greater generation quota. Under this mechanism, the proportion of H4/H5 items increases from 12.91%12.91\% to 25.25%25.25\% (Zhang et al., 2 Aug 2025).

Answer refinement is performed with DeepSeek-V3. Each synthesized QA pair is passed through a solvability filter that drops ill-posed questions. For solvable items, a step-by-step reasoning process regenerates a clean answer. After refinement, 60.7%60.7\%0 of answers differ from their initial versions, which indicates that answer post-processing is not merely cosmetic but materially alters a nontrivial share of outputs.

A recurrent misconception about synthetic QA generation is that diversification alone suffices to produce challenging data. The reported pipeline explicitly contradicts that view: broad STEM synthesis tends toward lower difficulty, and a separate difficulty-targeting mechanism is required to shift the hard-example distribution.

4. Mid-training formulation and integration

BoostQA is used in a mid-training stage situated between pre-training and fine-tuning. The reported rationale is to refine domain-specific knowledge acquisition while enhancing data quality (Zhang et al., 2 Aug 2025). Rather than replacing general text exposure, BoostQA is blended 1:1 with a high-quality general text corpus called KnowEdu, which is curated by QuRater for knowledge density and FineWeb-Edu for educational utility.

The mixture is defined as

60.7%60.7\%1

where 60.7%60.7\%2 and 60.7%60.7\%3.

The stated reason for this blend is to prevent the model from “forgetting” general language patterns. This is an important clarification: BoostQA is not presented as a standalone replacement for general pretraining text, but as one half of a deliberately balanced mid-training distribution.

In the main configuration, Llama-3 8B is continued from its 2T-token checkpoint for an additional 40B tokens over 60.7%60.7\%4. Training uses standard cross-entropy loss,

60.7%60.7\%5

the Adam optimizer, a learning-rate schedule from 60.7%60.7\%6 to 60.7%60.7\%7, and a global batch size of 960 (Zhang et al., 2 Aug 2025).

5. Reported evaluation results

The reported benchmark suite contains 12 evaluations: 5 knowledge-intensive benchmarks, 2 math benchmarks, 3 commonsense benchmarks, and 2 complex-reasoning benchmarks (Zhang et al., 2 Aug 2025). On this aggregate, BoostQA is reported to outperform all baselines and to achieve SOTA averages.

For the pre-training baseline, Llama-3 8B at 10T tokens attains MMLU 60.7%60.7\%8, CMMLU 60.7%60.7\%9, and AVG 33.8%33.8\%0. With mid-training on BoostQA + KnowEdu for 40B tokens, the same model reaches MMLU 33.8%33.8\%1, CMMLU 33.8%33.8\%2, and AVG 33.8%33.8\%3. The average gain on MMLU and CMMLU is reported as

33.8%33.8\%4

Across all 12 benchmarks, the reported absolute average improvement is 33.8%33.8\%5 (Zhang et al., 2 Aug 2025). The significance of these numbers lies in the fact that the gains are obtained in a mid-training regime rather than through a conventional post-training alignment phase. This suggests that a substantial portion of the improvement is attributed to knowledge acquisition and representation refinement at the corpus-and-training-distribution level.

BoostQA is reported to scale robustly with model size, data volume, and initial FLOPs (Zhang et al., 2 Aug 2025). For model size, final-checkpoint MMLU results are given as follows: at 1.7B parameters, KnowEdu yields 33.8%33.8\%6 and BoostQA yields 33.8%33.8\%7 for a gain of 33.8%33.8\%8; at 8B, KnowEdu yields 33.8%33.8\%9 and BoostQA yields 51.89%51.89\%0 for a gain of 51.89%51.89\%1; at 16B, KnowEdu yields 51.89%51.89\%2 and BoostQA yields 51.89%51.89\%3 for a gain of 51.89%51.89\%4.

For data volume on Llama-3 8B, 40B tokens produce MMLU 51.89%51.89\%5 with KnowEdu and 51.89%51.89\%6 with BoostQA, while 190B tokens produce MMLU 51.89%51.89\%7 with KnowEdu and 51.89%51.89\%8 with BoostQA. The reported improvement from 40B to 190B with BoostQA is 51.89%51.89\%9 on MMLU/CMMLU. For initial FLOPs, starting at 2T gives KnowEdu 10.87%10.87\%0 and BoostQA 10.87%10.87\%1, whereas starting at 10T gives KnowEdu 10.87%10.87\%2 and BoostQA 10.87%10.87\%3. The reported interpretation is that early-stage mid-training from the 2T checkpoint with BoostQA ultimately surpasses KnowEdu-only training even when the latter starts from the 10T checkpoint.

Throughout training, BoostQA yields the lowest smoothed language-model loss, and its stable, continually rising curve up to 40B tokens contrasts with the convergence observed in other blends. The work states that this is consistent with the well-known loss–performance correlation (Zhang et al., 2 Aug 2025). A plausible implication is that the benefits of BoostQA are not exhausted by short-horizon optimization and that the corpus continues to supply useful learning signals at scales where alternative mixtures begin to saturate.

Several clarifications follow directly from the reported setup. First, BoostQA targets STEM and high-difficulty weaknesses rather than uniformly all failure modes. Second, higher synthesis volume is not treated as sufficient in itself; the pipeline explicitly reweights generation toward H4/H5 and rejects items with estimated pass rate 10.87%10.87\%4 in the high-difficulty branch. Third, the method depends on retaining general text through KnowEdu, indicating that specialized QA synthesis is used as a complement to, not a substitute for, broad linguistic exposure. In the paper’s own summary, the result is a plug-and-play QA corpus whose disciplined seed curation, targeted annotation, dual-mode synthesis, and rigorous refinement remedy STEM and high-difficulty weaknesses, scale with model and data size, and deliver SOTA results in a mid-training regime (Zhang et al., 2 Aug 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to BoostQA.