BoostQA: 100B-Token QA Dataset
- BoostQA is a large-scale synthetic QA dataset designed to overcome scale, diversity, and difficulty limitations in traditional QA corpora.
- It employs a four-stage pipeline—seed curation, dual annotation, two-way synthesis, and answer refinement—to generate challenging, STEM-focused questions.
- When blended with KnowEdu in a mid-training regime, BoostQA achieves up to a 12.74% improvement on benchmarks, establishing state-of-the-art performance.
Searching arXiv for the specified paper to ground the article in the current record. BoostQA is a 100 billion-token, large-scale synthetic question-answering dataset introduced in "Large-Scale Diverse Synthesis for Mid-Training" (Zhang et al., 2 Aug 2025). It is designed to address three limitations in existing QA corpora for LLMs: insufficient scale, limited knowledge diversity—especially in cross-discipline settings—and the tendency of synthetic pipelines to generate lower-difficulty items. BoostQA is produced by a four-stage pipeline comprising seed curation, dual annotation, two-way synthesis, and answer refinement, and it is consumed in a mid-training stage situated between conventional pre-training and fine-tuning. In the reported setting, a 1:1 blend of BoostQA and KnowEdu yields an average improvement of on MMLU and CMMLU for Llama-3 8B and establishes SOTA average performance across 12 benchmarks (Zhang et al., 2 Aug 2025).
1. Problem setting and design objectives
The central motivation for BoostQA is the scarcity of high-quality, knowledge-intensive training data for LLMs. Traditional corpora are described as providing limited information, while previous work on synthesized QA is characterized as improving model performance but facing persistent problems in QA data scalability and knowledge diversity, particularly in cross-domain contexts (Zhang et al., 2 Aug 2025).
Three objectives organize the design. The first is scale: publicly available synthetic QA datasets rarely exceed a few tens of billions of tokens, whereas BoostQA expands the scale to 100B tokens. The second is diversity: traditional pipelines typically “rephrase and extend” narrow-domain passages, which yields limited cross-discipline reuse. The third is difficulty: probe experiments on Llama-3 8B checkpoints from 2T to 10T tokens, using a human-calibrated difficulty scorer, show steep degradation from easy to hard tiers, with accuracy falling from on H1 to on H5. This diagnosis motivates an explicit effort to amplify hard examples (Zhang et al., 2 Aug 2025).
In this formulation, BoostQA is not merely a larger synthetic corpus. It is structured as an intervention on identified weaknesses in STEM disciplines and high-difficulty data. This suggests that the dataset is intended to modify the composition of training signals rather than simply increase token count.
2. Corpus composition, coverage, and annotation
BoostQA draws its seeds from heterogeneous sources: public QA datasets such as MathQA, high-quality textbooks, and specialized web crawls such as MegaMath-Web (Zhang et al., 2 Aug 2025). Before synthesis, the pipeline performs decontamination using exact 10-gram matching plus embedding-based similarity filtering against all evaluation benchmarks. This places benchmark isolation at the corpus-construction stage rather than treating it as a downstream evaluation concern.
Two annotation systems drive both synthesis and analysis. A discipline classifier with 62 categories is implemented via Qwen2.5-7B-Instruct and is distilled from DeepSeek-R1. A separate difficulty scorer assigns H1–H5 tiers and is calibrated to human pass rates under one-hour tests by top-100 QS students. These labels determine synthesis direction and are also used for distributional analyses (Zhang et al., 2 Aug 2025).
| Attribute | Reported value |
|---|---|
| Total size | 100 billion tokens |
| Discipline coverage | 62 primary academic disciplines |
| Top discipline shares | Mathematics 51.89%, Computer Science 10.87%, Clinical Medicine 6.25% |
| Grade levels in synthetic data | High School, College, Graduate |
| Difficulty distribution | H1/H2 = 56.86%, H3 = 15.59%, H4/H5 = 27.55% |
The discipline profile is explicitly skewed, with Mathematics comprising of the corpus, followed by Computer Science at and Clinical Medicine at . The synthetic data is organized by grade levels—High School, College, and Graduate—but the reported pipeline also observes a down-shift in actual alignment and compensates by prompting one tier higher. A plausible implication is that nominal prompt role and realized item difficulty are not identical, which is why explicit post hoc difficulty control remains necessary.
3. Four-stage synthesis and refinement pipeline
The pipeline begins with seed curation, followed by discipline and difficulty annotation, then two-way synthesis with DeepSeek-R1, and finally answer refinement with DeepSeek-V3 (Zhang et al., 2 Aug 2025). The four-stage decomposition is central to the identity of BoostQA and distinguishes it from approaches that only extend pre-existing QA pairs.
The first synthesis branch is STEM-focused multi-grade synthesis. It uses role prompts corresponding to “High School,” “College,” and “Graduate” educators. Prompting enforces format constraints; for example, multiple-choice items must contain four distractors, while essay questions require a self-contained prompt plus an explanatory solution. The output size is configurable, typically per seed, and the seed set is over-sampled toward STEM. This branch yields broad coverage but tends toward lower-difficulty items.
The second branch is high-difficulty synthesis, described as a Difficulty Booster. Here the role is fixed to “Graduate,” and an integrated difficulty validator rejects any item with estimated pass rate . Sampling weights are re-balanced so that seeds labeled H4/H5 receive greater generation quota. Under this mechanism, the proportion of H4/H5 items increases from to (Zhang et al., 2 Aug 2025).
Answer refinement is performed with DeepSeek-V3. Each synthesized QA pair is passed through a solvability filter that drops ill-posed questions. For solvable items, a step-by-step reasoning process regenerates a clean answer. After refinement, 0 of answers differ from their initial versions, which indicates that answer post-processing is not merely cosmetic but materially alters a nontrivial share of outputs.
A recurrent misconception about synthetic QA generation is that diversification alone suffices to produce challenging data. The reported pipeline explicitly contradicts that view: broad STEM synthesis tends toward lower difficulty, and a separate difficulty-targeting mechanism is required to shift the hard-example distribution.
4. Mid-training formulation and integration
BoostQA is used in a mid-training stage situated between pre-training and fine-tuning. The reported rationale is to refine domain-specific knowledge acquisition while enhancing data quality (Zhang et al., 2 Aug 2025). Rather than replacing general text exposure, BoostQA is blended 1:1 with a high-quality general text corpus called KnowEdu, which is curated by QuRater for knowledge density and FineWeb-Edu for educational utility.
The mixture is defined as
1
where 2 and 3.
The stated reason for this blend is to prevent the model from “forgetting” general language patterns. This is an important clarification: BoostQA is not presented as a standalone replacement for general pretraining text, but as one half of a deliberately balanced mid-training distribution.
In the main configuration, Llama-3 8B is continued from its 2T-token checkpoint for an additional 40B tokens over 4. Training uses standard cross-entropy loss,
5
the Adam optimizer, a learning-rate schedule from 6 to 7, and a global batch size of 960 (Zhang et al., 2 Aug 2025).
5. Reported evaluation results
The reported benchmark suite contains 12 evaluations: 5 knowledge-intensive benchmarks, 2 math benchmarks, 3 commonsense benchmarks, and 2 complex-reasoning benchmarks (Zhang et al., 2 Aug 2025). On this aggregate, BoostQA is reported to outperform all baselines and to achieve SOTA averages.
For the pre-training baseline, Llama-3 8B at 10T tokens attains MMLU 8, CMMLU 9, and AVG 0. With mid-training on BoostQA + KnowEdu for 40B tokens, the same model reaches MMLU 1, CMMLU 2, and AVG 3. The average gain on MMLU and CMMLU is reported as
4
Across all 12 benchmarks, the reported absolute average improvement is 5 (Zhang et al., 2 Aug 2025). The significance of these numbers lies in the fact that the gains are obtained in a mid-training regime rather than through a conventional post-training alignment phase. This suggests that a substantial portion of the improvement is attributed to knowledge acquisition and representation refinement at the corpus-and-training-distribution level.
6. Scaling behavior, loss trends, and interpretive clarifications
BoostQA is reported to scale robustly with model size, data volume, and initial FLOPs (Zhang et al., 2 Aug 2025). For model size, final-checkpoint MMLU results are given as follows: at 1.7B parameters, KnowEdu yields 6 and BoostQA yields 7 for a gain of 8; at 8B, KnowEdu yields 9 and BoostQA yields 0 for a gain of 1; at 16B, KnowEdu yields 2 and BoostQA yields 3 for a gain of 4.
For data volume on Llama-3 8B, 40B tokens produce MMLU 5 with KnowEdu and 6 with BoostQA, while 190B tokens produce MMLU 7 with KnowEdu and 8 with BoostQA. The reported improvement from 40B to 190B with BoostQA is 9 on MMLU/CMMLU. For initial FLOPs, starting at 2T gives KnowEdu 0 and BoostQA 1, whereas starting at 10T gives KnowEdu 2 and BoostQA 3. The reported interpretation is that early-stage mid-training from the 2T checkpoint with BoostQA ultimately surpasses KnowEdu-only training even when the latter starts from the 10T checkpoint.
Throughout training, BoostQA yields the lowest smoothed language-model loss, and its stable, continually rising curve up to 40B tokens contrasts with the convergence observed in other blends. The work states that this is consistent with the well-known loss–performance correlation (Zhang et al., 2 Aug 2025). A plausible implication is that the benefits of BoostQA are not exhausted by short-horizon optimization and that the corpus continues to supply useful learning signals at scales where alternative mixtures begin to saturate.
Several clarifications follow directly from the reported setup. First, BoostQA targets STEM and high-difficulty weaknesses rather than uniformly all failure modes. Second, higher synthesis volume is not treated as sufficient in itself; the pipeline explicitly reweights generation toward H4/H5 and rejects items with estimated pass rate 4 in the high-difficulty branch. Third, the method depends on retaining general text through KnowEdu, indicating that specialized QA synthesis is used as a complement to, not a substitute for, broad linguistic exposure. In the paper’s own summary, the result is a plug-and-play QA corpus whose disciplined seed curation, targeted annotation, dual-mode synthesis, and rigorous refinement remedy STEM and high-difficulty weaknesses, scale with model and data size, and deliver SOTA results in a mid-training regime (Zhang et al., 2 Aug 2025).