Papers
Topics
Authors
Recent
Search
2000 character limit reached

Darwin-Science: Data Darwinism for STEM AI

Updated 3 July 2026
  • Darwin-Science is a framework that applies Data Darwinism principles to transform raw scientific texts into high-information corpora using a ten-level taxonomy.
  • It employs advanced generative refinement (L4) and cognitive completion (L5) to enhance semantic clarity and reconstruct explicit scientific reasoning.
  • Empirical benchmarks demonstrate significant gains on in-domain and general tasks, validating the approach’s impact on scientific machine learning.

Darwin-Science (“Data Darwinism”)

Darwin-Science refers to a framework, dataset, and methodology arising from the application of the principles of “Data Darwinism” to scientific literature and data-centric AI. It is an instantiation of a ten-level taxonomy for data-model co-evolution, aiming to unlock the latent value of complex scientific texts and catalyze advancements in scientific machine learning through systematic data processing and model pre-training. The Darwin-Science corpus—comprising ∼900B processed tokens—serves as both proof of principle and practical resource for foundation model (FM) development in STEM domains. The approach demonstrates that naïvely scaling raw scientific data does not improve scientific reasoning in models; only through targeted generative refinement and cognitive completion (with advanced LLMs) can non-trivial performance gains be realized in downstream scientific benchmarks (Qin et al., 8 Feb 2026).

1. Data Darwinism Taxonomy: Conceptual Foundations

The Data Darwinism taxonomy defines a hierarchical, evolutionary sequence of data refinement and augmentation, transforming raw data into high-density, model-friendly corpora through a series of increasingly complex processing stages (Qin et al., 8 Feb 2026):

Level Summary Key Operations
L0 Data Acquisition Raw collection (HTML, PDF, scans, AV)
L1 Format Normalization OCR, parsing, transcript generation
L2 Rule-Based Filtering Deduplication, heuristics (length, language, quality)
L3 Lightweight Model Filtering Small-LM classification for discipline/type/quality
L4 Generative Refinement LLM-based cleaning (removing garble, references, artifacts)
L5 Cognitive Completion LLM-driven reasoning chain explication, inlined definitions
L6–L9 Contextual/World Synthesis (future) Retrieval, simulation, environment/ecosystem synthesis

Levels L0–L3 emphasize selection and preservation, while L4–L5 explicitly introduce structure, clarity, and accessibility through LLM-guided intervention. L6–L9, not yet fully realized, address dynamic knowledge integration and synthetic simulation. The evolutionary metaphor is articulated both in terms of volume–information density trade-offs and in practical data-model feedback loops (Qin et al., 8 Feb 2026).

2. Darwin-Science Corpus Construction and Processing Pipeline

Darwin-Science implements the L0–L5 pipeline concretely for scientific books and papers. The pipeline is as follows (Qin et al., 8 Feb 2026):

  • L0: Acquire ∼50M documents (public scans of academic books; arXiv, PubMed Central, S2ORC full text), totaling 601B raw tokens.
  • L1: Format normalization via vision-language OCR (olmOCR-7B-0225-preview), segmenting documents into 1,024-character text chunks.
  • L2: Rule-based filtering to remove near-duplicates (MinHash+LSH), enforce language and quality thresholds, and discard artifacts.
  • L3: Lightweight semantic filtering with EAI-Distill-0.5B to classify for relevance (12-dim domain/quality/type labels); removal of non-educational content.
  • L4: Generative refinement using large LLMs (GPT-OSS-120B); rules enforce faithfulness: removal of references, metadata, OCR error correction, format repair, and structural standardization.
  • L5: Cognitive completion, applied to papers, using Qwen3-235B; the model is prompted to reconstruct reasoning chains, define terms in context, and supply analogies, stepwise derivations, and pedagogical interleaving, all under strict factual integrity.

The resulting corpus contains ~906.5B tokens, with books (2.98M docs, L4 only, avg 84k tokens/doc) and papers (47.8M docs; L4 avg 8.1k tokens/doc; L5 avg 20.5k tokens/doc). A rigorous contamination filter (exact 20-gram match) is applied to remove evaluation benchmark overlap, ensuring “clean-room” attribution (Qin et al., 8 Feb 2026).

3. Mechanisms of Generative Refinement and Cognitive Completion

Levels L4 and L5 are operationalized by prompt engineering and output constraints on advanced instruction-tuned LLMs (Qin et al., 8 Feb 2026):

  • L4 (Generative Refinement): Removes low-information artifacts (metadata, ToC, junk), repairs format splits, and restores semantic continuity in 1,024-token sliding windows. Prompts codify allowed deletions/modifications; failed chunks are flagged for reprocessing.
  • L5 (Cognitive Completion): Expands terse or implicit text into explicit reasoning. The model is prompted as a “master science communicator” bound by mandates: maintain 1:1 structural mapping, preserve formal content, explicate proofs or derivations, and bridge to intuition with analogies. Stepwise exposition is enforced, but the model is forbidden to introduce ungrounded (“hallucinated”) content.

Empirical examples demonstrate L5 rewriting of theorem statements to include expanded definitions, test-function arguments, and contextual explanations. L5 also restructures dense paragraphs into stepwise guides with explicit physical units and procedural clarity.

4. Quantitative Evaluation: Metrics, Ablation, and Learnability Gap

Performance is assessed via continued pre-training (CPT) of from-scratch, contamination-free daVinci-origin-3B/7B models on the Darwin-Science corpus versus a carefully matched baseline (no scientific content) (Qin et al., 8 Feb 2026). Metrics include average accuracy across general, scientific, and in-domain benchmarks (e.g., BBH, ARC-Easy/Challenge, MMLU, GSM-8K, MATH, SuperGPQA, SciBench, Darwin-Science-Eval).

Key results:

  • CPT Gains: +2.12 (3B), +2.95 (7B) average benchmark improvement post-CPT on Darwin-Science.
  • In-Domain Gains: +5.60 (3B), +8.40 (7B) on domain-aligned tasks (Darwin-Science-Eval).
  • Ablation: Raw L0–L3 data gives near-zero gain; L4 adds +0.38; L5 lifts to +1.36 (total gain) [ΔP_{L4} ≈ 0.38; ΔP_{L5} ≈ 0.98].
  • Learnability Gap: Without L4/L5, additional scientific data supplies negligible marginal value (ΔP ≈ 0), demonstrating that unprocessed complexity is not sufficient for model learning.
  • Optimal Mixture: ~50% domain data maximizes gains on general benchmarks; increasing the scientific fraction further boosts in-domain performance.

Empirical formulas: ΔP=PDarwin−PBaseline,ΔP‾=1N∑i=1N(PiDarwin−PiBaseline)\Delta P = P_{\rm Darwin} - P_{\rm Baseline}, \quad \overline{\Delta P} = \frac{1}{N}\sum_{i=1}^N (P_i^{\rm Darwin} - P_i^{\rm Baseline})

Glearn=PRaw−PBaseline≈0\mathcal{G}_{\rm learn} = P_{\rm Raw} - P_{\rm Baseline} \approx 0

5. Methodological Controls, Limitations, and Future Trajectories

Strict contamination controls are enforced: base models are trained on scientific-free data; Darwin-Science is decontaminated via benchmark n-gram match eliminating ~0.03% overlap (Qin et al., 8 Feb 2026). This yields closed attribution loops for data impact.

Limitations include:

  • L6–L9 (retrieval, synthesis environments, simulation) remain prospective.
  • Only text modality; multimodal data (figures, code, labs) is not yet integrated.
  • Teacher LLM selection is fixed; full cost-benefit landscape for generative refinement is not exhaustively mapped.
  • Learnability gap is empirically quantified, but formal theoretical metrics warrant further development.

Future directions center on automated co-evolutionary L0–L9 pipelines, multimodal data integration, finer-grained assessment of data utility, and formalization of metrics like learnability prediction.

6. Scientific and Practical Significance

Darwin-Science exemplifies a systematic approach to knowledge distillation from complex, high-entropy scientific data. Empirical results demonstrate that higher-level model-driven processing (L4/L5) is essential to unlock domain-specific performance in FMs, revealing that data and models should co-evolve rather than scale naïvely (Qin et al., 8 Feb 2026). The public release of Darwin-Science and daVinci-origin models offers reproducible baselines for the community, accelerating principled data-centric research in scientific machine learning.

This framework operationalizes the evolutionary analogy from “Universal Darwinism” to data curation and model development: only through cycles of selection (filtering), variation (generative rewriting), and retention (canonicalization) do datasets attain the density and structure required for effective learning and downstream scientific reasoning.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Darwin-Science.