---
title: 'FinePhrase: Synthetic Data & ESL Assessment'
url: https://www.emergentmind.com/topics/finephrase
type: topic
---

# FinePhrase: Synthetic Data & ESL Assessment

FinePhrase refers to two distinct yet influential methodologies: (1) an open framework and dataset for high-quality synthetic language model pretraining via systematic web text rephrasing, and (2) an approach for automated phrase-break assessment in ESL learner speech using pre-trained language models (PLMs) and large language models (LLMs). Below, both strands are delineated, with primary emphasis on the large-scale synthetic data paradigm established in "How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data" [2604.13977]. The ESL assessment framework as described in "Assessing Phrase Break of ESL Speech with Pre-trained Language Models and Large Language Models" [2306.04980] is also summarized for completeness.

## 1. Motivation and Foundational Objectives

The FinePhrase synthetic data corpus addresses a specific bottleneck in language model pretraining: the stagnation in the availability of novel, high-quality web-scale text. While early synthetic datasets such as WRAP, Nemotron-CC, and REWIRE provided valuable tokens using various generators and prompt templates, they did not systematically compare prompt structure, generator scaling, or data mixing. FinePhrase systematically varies these components to determine how rephrasing strategy (prompt format), generator scale, and source/mix-in data impact downstream benchmark performance, cost-efficiency, and practical usability. The explicit goals are to (1) identify optimal synthetic data formats, (2) quantify necessary generator model size, (3) evaluate source and mix-in data effects, and (4) deliver an open, large-scale (486B-token) dataset that substantially outperforms prior synthetic and real-data baselines at up to 30× lower compute cost [2604.13977].

## 2. Rephrasing Strategies and Prompt Engineering

Eight existing rephrasing prompt types (Diverse QA Pairs, Guided Rewrite, Summarize, Continue, Distill, Wikipedia-style paraphrasing, Knowledge List, Extract Knowledge) are compared with four novel pedagogical formats, each transforming raw text into formats that are more structured, discrete, and information-dense:

| Format    | Description                                                             | Example Task                   |
|-----------|------------------------------------------------------------------------|--------------------------------|
| math      | Create a mathematical word problem; include step-by-step solution       | Problem generation             |
| faq       | Build a standalone FAQ; questions ordered from foundational to advanced | Knowledge distillation         |
| table     | Structure key information in a table; add one QA pair per table         | Entity and fact organization   |
| tutorial  | Rewrite as a procedural tutorial with steps or bullets                  | Skill and procedure instruction|

In mathematical terms: given a document $D \in \mathcal{D}_s$, generate a completion $C = G(\mathcal{P}(D))$ where $\mathcal{P}$ is the prompt transformation, and $G$ is the generator model [2604.13977].

Empirical findings demonstrate that structured pedagogical prompts—math, table, FAQ, tutorial—consistently outperform both ad-hoc and prior synthetic approaches across most evaluated downstream tasks.

## 3. Generator Model Scaling and Family Effects

FinePhrase evaluates both Gemma 3 (270M, 1B, 4B, 12B, 27B) and SmolLM2 (135M, 360M, 1.7B) model families for synthetic data generation efficiency and efficacy. Downstream LM performance, measured by macro average scores over 12 benchmarks, saturates at approximately 1B parameters across rephrasing formats. For instance, Gemma 3 at 1B parameters achieves macro-avg 15.31; at 27B, it achieves 14.76, while requiring up to 10× greater compute. Among six 1B-scale generator families, SmolLM2 delivers the best results (16.55 vs. DCLM’s 13.77), attributed to superior instruction-tuning on rewrite tasks. This establishes strong diminishing returns for large generator scaling and favors parameter- and energy-efficient approaches [2604.13977].

## 4. Source Data and Mix-In Strategy

FinePhrase distinguishes source data ($\mathcal{D}_s$; the input to be rephrased) from mix-in data ($\mathcal{D}_m$; original real tokens included for pretraining diversity). Exclusive synthetic pretraining degrades natural language understanding (NLU) capabilities; mixing synthetic and high-quality web data at a 50/50 ratio restores and enhances NLU performance (as indicated by improvements on WinoGrande, HellaSwag, etc., in Table 8). Best results arise from mixing FineWeb-HQ or DCLM as the real-data counterpart. Notably, even when the source data ($\mathcal{D}_s$) is lower quality, competitive performance is recoverable by pairing with robust mix-in ($\mathcal{D}_m$), confirming the up-cycling effect of high-fidelity prompting on noisy input [2604.13977].

## 5. Corpus Composition, Scale, and Key Benchmarks

FinePhrase synthesizes a 486B-token dataset by applying four pedagogical formats to a filtered FineWeb web crawl, yielding 1.35B structured samples. Each format (math, faq, table, tutorial) contributes roughly 121.5B tokens. Models pretrained from scratch on 21B tokens of 50/50 synthetic+FineWeb-HQ data (using 1.2B-param Qwen 2 LMs) are evaluated via 3-shot prompting across 12 benchmarks (ARC, MMLU, SQuAD v2, DROP, GSM8K, OpenBookQA, XCSQA, WinoGrande, PIQA, HellaSwag, WikiTableQuestions, TriviaQA).

Empirical results: FinePhrase (table format) delivers macro-avg 17.18, compared to 13.77 for DCLM and 13.54 for Nemotron-HQ-Synth. Math and FAQ formats also achieve macro-avg 16.97 and 16.18, respectively. Structured formats excel particularly on factual knowledge and reading comprehension tasks (e.g., +18.72 on SQuAD v2 over baselines), but less so on NLU-oriented tasks unless real-data mixing is employed [2604.13977].

## 6. Efficiency and Cost Analysis

FinePhrase's use of speculative decoding (9.2k tokens/sec per H100 GPU) and efficient model design yields mean throughput of 33.1M tokens/GPU-hr. Generation of all 486B tokens required approximately 14.7K GPU-hr (612 GPU-days). In comparative terms, FinePhrase achieves a 30× efficiency gain relative to REWIRE, and a 13× gain relative to Cosmopedia, in tokens generated per GPU-hr, as detailed below:

| Dataset      | Generator         | Tokens | GPU-hrs | Tokens/GPU-hr |
|--------------|-------------------|--------|---------|--------------|
| Cosmopedia   | Mixtral 8×7B      | 25B    | >10K    | <2.5M        |
| REWIRE       | Llama 3.3 70B     | 400B   | ~352K   | ~1.1M        |
| FinePhrase   | SmolLM2 1.7B      | 486B   | ~14.7K  | ~33.1M       |

[2604.13977]

## 7. Design Principles and Recommendations

Optimal synthetic pretraining corpus synthesis per FinePhrase evidence requires:

- Prioritization of pedagogical prompt design (math, table, FAQ, tutorial) over aggressive generator scaling.
- Selection of ∼1B-parameter, instruction-tuned generators for efficiency without loss of data quality.
- Systematic mixing of synthetic and real, high-quality corpora to maintain language diversity and NLU robustness.
- Employment of robust mix-in data (FineWeb-HQ, DCLM) when working from low-quality sources to exploit up-cycling effects.
- Encouragement of output diversity via generator family selection (SmolLM2 preferred) to mitigate template collapse.
- Adoption of advanced inference frameworks (e.g., speculative decoding, vLLM) for maximizing hardware productivity.
- Commitment to openness: all prompts, generation frameworks, and data assets are released for reproducibility and further research [2604.13977].

## 8. FinePhrase for ESL Phrase-Break Assessment

In the distinct domain of ESL speech assessment, FinePhrase denotes a PLM/LLM pipeline for evaluating phrase-break quality in learner speech [2306.04980]. The system consists of:

- Forced alignment of ESL speech and transcript to extract interword pause durations, discretized into break tokens ($b_i \in \{\mathtt{br}_0, \mathtt{br}_1, \mathtt{br}_2, \mathtt{br}_3\}$).
- Sequential encoding $T = \{w_0, b_0, w_1, \ldots, w_n\}$ for input to PLMs.
- Pre-training via replaced-break-token detection on corrupted TTS data, with binary cross-entropy loss and 3:1 corrupted:original data composition.
- Task-specific fine-tuning: overall assessment (sequence classification, ranks 1–3), and fine-grained labeling of break tokens (label 1–3).
- LLM-based evaluation (ChatGPT) using prompt engineering; zero-shot and few-shot prompts improve performance (overall F1_macro from 40.6 to 47.3).
- Superiority of Break-BERT over non-PLM and vanilla BERT (boosting macro-F1 in overall assessment from 41% to 52%, and fine-grained token classification macro-F1 from ~40% to 44.3%).

This pipeline establishes a multi-pronged methodology combining self-supervised TTS pre-training, supervised fine-tuning, and LLM-based analysis in a single, extensible framework for spoken language assessment.

## References

- "How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data" [2604.13977]
- "Assessing Phrase Break of ESL Speech with Pre-trained Language Models and Large Language Models" [2306.04980]

Source: https://www.emergentmind.com/topics/finephrase