---
title: 'LingVarBench: Synthetic NER Benchmark for Healthcare Calls'
url: https://www.emergentmind.com/topics/lingvarbench
type: topic
---

# LingVarBench: Synthetic NER Benchmark for Healthcare Calls

LingVarBench is a synthetic benchmark and pipeline for training and evaluating large language model–based named entity recognition and structured extraction systems on phone-call style conversational speech, particularly in healthcare settings such as patient–provider calls involving ZIP codes, dates of birth, and names. It addresses the problem that labeling real phone transcripts is expensive, privacy-sensitive, and operationally constrained, while conventional NER benchmarks do not reflect the disfluencies, interruptions, informal syntax, and formatting variability of conversational speech. In this formulation, LingVarBench is both a benchmark and a method: it generates structured synthetic utterances, validates recoverability of the underlying entities, and uses the resulting data to optimize extraction prompts that transfer to real customer transcripts [2508.15801].

## 1. Problem formulation and scope

LingVarBench is motivated by the need to extract structured information from spontaneous speech in healthcare voice AI systems. The target data regime is phone call transcripts containing disfluencies, self-corrections, hesitations and pauses, filler words, variable formatting, background noise and ASR errors in real data, and domain-specific healthcare terminology and emotional content. The benchmark focuses on fields that are operationally important for authentication and downstream workflow, including ZIP code, Date of Birth, and Name [2508.15801].

The central obstacle is the scarcity of large, high-quality, HIPAA-compliant labeled datasets for conversational healthcare speech. The reported constraints are concrete: phone call transcript labeling is “approximately 2 USD per minute,” and manual annotation requires “3 hours of expert time per hour of audio.” The same setting also introduces privacy and governance burdens, including PHI handling, consent requirements, BAAs, and privacy review. LingVarBench is designed as a way to bypass these barriers by replacing direct annotation of real transcripts with validated synthetic conversational data and automated prompt optimization [2508.15801].

This design distinguishes LingVarBench from standard written-text NER benchmarks such as CoNLL-2003, OntoNotes, and clinical NER over written notes. Those resources do not capture phone-call style conversational phenomena and are not constructed under the same regulatory constraints. A further distinction is that LingVarBench is domain-specific rather than multilingual in the sense used by broad equity-oriented benchmarks such as GlobalBench, which tracks 966 datasets in 190 languages and evaluates utility and equity across languages [2305.14716]. LingVarBench instead concentrates on structured extraction under conversational variation in a healthcare workflow.

## 2. Generative architecture and validation loop

LingVarBench is organized as a three-stage generative pipeline followed by prompt optimization. The stages are a Value Generator \(V\), a Transcript Generator \(T\), and a Validator \(\mathcal{C}\). The Value Generator produces plausible canonical structured values for a field description; the Transcript Generator transforms those values into natural conversational utterances under specified variation types; and the Validator retains only utterances from which the target value can be recovered. The validated corpus is then used inside DSPy with SIMBA to optimize extraction prompts [2508.15801].

Formally, the paper defines entity types \(e \in \mathcal{E}\), field descriptions \(d \in \mathcal{D}\), variation types \(v \in \mathcal{V}\), and utterances \(u \in \mathcal{U}\), with component mappings

\[
V: \mathcal{D} \rightarrow \mathbb{V}, \qquad
T: (\mathbb{V}, \mathcal{V}) \rightarrow \mathcal{U}, \qquad
\mathcal{C}: \mathcal{U} \rightarrow (\mathcal{U}, e).
\]

The composite generator is

\[
G(d, v) = \mathcal{C}\left(T\left(V(d),\, v\right)\right), \quad
G : (\mathcal{D}, \mathcal{V}) \rightarrow (\mathcal{U}, e).
\]

This means that, for each field description and variation type, LingVarBench produces a validated pair \((u,e)\): a transcript containing a recoverable entity instance [2508.15801].

A concise view of the workflow is as follows:

| Stage | Function | Output |
|---|---|---|
| Value Generator \(V\) | Samples plausible canonical values from a field description | Structured values such as “91101” or “12-02-1947” |
| Transcript Generator \(T\) | Verbalizes a canonical value under specified variation types | Spoken-style transcripts with variation tags |
| Validator \(\mathcal{C}\) | Tests whether the value is recoverable from the transcript | Filtered transcript–entity pairs |
| DSPy + SIMBA | Optimizes extraction prompts on validated synthetic data | Entity-specific prompts for real and synthetic evaluation |

The validator is a critical quality-control mechanism. Generated utterances are not accepted by default; they are checked by a separate LLM acting as a binary extractor or verifier. If the transcript contains the value, even with corrections, validation returns `true`; if it is vague or does not contain the value, it returns `false`. For dates, the validation logic explicitly accepts both continuous digits and spoken sequences. Invalid transcripts are discarded, and generation may be recursively re-invoked to fill the quota for a value–variation pairing. The benchmark is therefore not merely synthetic generation but “generate, verify, filter” [2508.15801].

## 3. Linguistic variation taxonomy and dataset construction

A defining feature of LingVarBench is its explicit taxonomy of linguistic variation types. Each field description comprises a field name, a data type, and a natural-language description specifying format or constraints. The transcript generation stage then conditions on one or more variation types that model how structured values are realized in speech [2508.15801].

The variation taxonomy has two layers. First, there are general variations that can apply to any entity, including `filler_words`, `hesitation`, `correction` or `self-correction`, `repetition`, `pause`, `formal`, `casual`, `polite`, `confident`, `uncertain`, `rushed`, `careful`, `confirmation`, and `direct_and_simple`. Second, there are entity-specific variations. For ZIP code, these include `digit_by_digit`, `grouped_two`, `grouped_three`, `hundred`, `mixed_grouping`, and `reversed`, along with combinations involving hesitations, fillers, and corrections. For DOB, the taxonomy includes `date_as_4_digits`, `date_as_5_digits`, `date_as_6_digits`, `date_as_8_digits`, `spoken_date_4_digits` through `spoken_date_8_digits`, `spoken_month_day_year`, `mixed_spoken_and_digits`, `filler_or_correction`, and `casual_or_polite_digits`. For Name, the benchmark includes `name_with_last`, `name_with_prefix`, `name_reverse_order`, `name_with_title`, `name_with_middle`, `name_with_suffix`, `name_with_initials`, `name_with_correction`, `name_partial_spelling`, `name_with_apostrophe`, `name_hyphenated`, and `nickname` [2508.15801].

The transcript generator pairs canonical values with these variation types, generates utterances in JSON form, and tracks tag frequencies. To avoid concentration on easy or highly frequent realizations, the system re-prompts for underrepresented value–variation pairs until a balanced distribution is reached. The benchmark description states that this yields “uniform coverage over linguistic variation types while preserving diversity at the utterance level” [2508.15801].

After generation and validation, the corpus is split into canonical train, validation, and test sets at 70% / 15% / 15%. The reported summary statistics show that the benchmark is substantial even in its initial scope:

| Entity | GPT-4 synthetic set | Gemini 2.0 Flash synthetic set |
|---|---|---|
| ZIP code | 5,635 samples; 33 variation tags; 10 unique canonical values; avg length ~41 tokens | 6,467 samples; 31 tags; 12 values |
| Name | 2,550 samples; 46 tags; 12 values | 6,631 samples; 46 tags; 12 values |
| DOB | 1,055 samples; 36 tags; 7 values | 682 samples; 33 tags; 5 values |

The corpus structure makes LingVarBench a benchmark rather than only a generation procedure: each sample has a transcript, a ground-truth canonical value, and associated variation tags. The benchmark is therefore suitable for both supervised prompt optimization and controlled robustness analysis across conversational phenomena [2508.15801].

## 4. DSPy, SIMBA, and automated prompt optimization

The synthetic corpus is used as training signal for automated prompt optimization rather than as an end in itself. LingVarBench integrates DSPy, described as a framework for programming LLM pipelines declaratively and automatically optimizing prompts and hyperparameters, with SIMBA, the Stochastic Introspective Mini-Batch Ascent optimizer. The explicit objective is to eliminate manual prompt engineering while learning prompts that are robust to conversational variation [2508.15801].

For each entity type, the procedure is: generate synthetic labeled transcripts; define a base extraction module in DSPy with general and field-specific instructions; optimize the prompt on synthetic training and validation data; and then evaluate the resulting prompt on both synthetic held-out data and real phone transcripts. The benchmark description gives representative field-specific instructions such as “Extract exactly 5 numeric digits,” “Do not interpret ZIP codes as dates,” and “Return the raw 5-digit number only” [2508.15801].

The extractor is treated as a prompt-parameterized function \(f_\theta(u)\). The optimization target is stated either as exact-match accuracy,

\[
\frac{1}{N} \sum_{i=1}^{N} \mathbb{1}[f_{\theta}(u_i) = y_i],
\]

or equivalently as F1,

\[
\text{F1}(\theta) = 2 \times \frac{\text{P}(\theta) \cdot \text{R}(\theta)}{\text{P}(\theta) + \text{R}(\theta)},
\qquad
\text{P} = \frac{TP}{TP + FP}, \quad
\text{R} = \frac{TP}{TP + FN}.
\]

SIMBA then performs mini-batch, introspective updates over candidate prompts, evaluating rephrasings and constraints on synthetic data and ascending toward prompts with better extraction performance [2508.15801].

This optimization strategy is central to the benchmark’s claim of practicality. The prompt search uses only synthetic validated transcripts and therefore avoids using PHI-containing real transcripts during optimization. In that sense, LingVarBench is not merely a synthetic dataset for model pretraining; it is a closed-loop system for generating task-specific training data, validating semantic faithfulness, and using the result to optimize LLM extraction behavior under privacy constraints [2508.15801].

## 5. Experimental setup and empirical results

Generation, validation, and extraction are instantiated with GPT-4, Gemini 2.0 Flash, and Gemini 2.5 Pro, using standardized prompt formats and deterministic decoding with temperature \(= 0\) and top-p \(= 1.0\). Similarity analysis uses `text-embedding-3-large` and `gemini-embedding-001`. Evaluation is conducted on both the synthetic LingVarBench splits and a proprietary real dataset from Infinitus Systems Inc. consisting of real patient-facing phone calls reviewed in a human-in-the-loop process for ZIP code, Name, and DOB ground truth [2508.15801].

The baseline comparison includes three prompt conditions: a zero-shot prompt, a human prompt iteratively refined with real data, and a LingVarBench-optimized prompt obtained through DSPy/SIMBA on synthetic data alone. On synthetic data, the optimized prompts achieve high accuracy: ZIP code reaches approximately 97% for GPT-4 and approximately 89% for Gemini 2.0 Flash; Name reaches approximately 96–97% for GPT-4 and approximately 79% for Gemini 2.0 Flash; DOB reaches approximately 98% for GPT-4, while Gemini 2.5 Pro reaches 100% on validation and approximately 98% on test [2508.15801].

The principal reported results concern transfer to real transcripts. For ZIP code, zero-shot performance is 88.62% accuracy for GPT-4 and 89.37% for Gemini 2.5 Pro; the best human prompt reaches 89.76% for Gemini 2.0 Flash; and LingVarBench-optimized prompting reaches 95.07% accuracy and F1 94.61 for Gemini 2.5 Pro, with GPT-4 at 89.79% accuracy and F1 94.66. For Name, zero-shot ranges from 47.87% accuracy for Gemini 2.0 Flash to 78.19% for GPT-4, whereas LingVarBench-optimized prompts reach 90.18% accuracy and F1 94.84 for both GPT-4 and Gemini 2.0 Flash. For DOB, zero-shot is 74.52% for GPT-4 and 77.04% for Gemini 2.5 Pro; human prompting reaches 82.62% for Gemini 2.0 Flash; and LingVarBench-optimized prompting reaches 80.80% accuracy and F1 89.38 for Gemini 2.5 Pro, and 80.66% accuracy and F1 89.30 for Gemini 2.0 Flash [2508.15801].

The paper summarizes these improvements as up to 95 percent accuracy for numeric fields versus 88–89 percent zero-shot, around 90 percent for names versus 47–79 percent zero-shot, and over 80 percent for dates versus 72–77 percent zero-shot. The resulting interpretation is narrow but strong: optimized synthetic-data prompts consistently outperform zero-shot prompts, match or surpass human-tuned prompts for ZIP and Name, and improve over zero-shot for DOB, although human-tuned prompts may remain slightly stronger in some DOB settings [2508.15801].

A further evaluation axis is synthetic-to-real similarity. Using cosine similarity over embeddings, when variation types match, `text-embedding-3-large` yields 0.81 ± 0.13 for ZIP, 0.81 ± 0.15 for Name, and 0.62 ± 0.14 for DOB; `gemini-embedding-001` yields 0.91 ± 0.07 for ZIP, 0.92 ± 0.08 for Name, and 0.83 ± 0.08 for DOB. Similarity drops when variation tags diverge, which the paper uses to argue that explicit variation control makes synthetic utterances closer to real conversational patterns [2508.15801].

## 6. Interpretation, limitations, and relation to adjacent benchmarks

The main analytical claim of LingVarBench is that controlled linguistic diversity plus automated prompt optimization can produce robust extractors for real conversational speech without using real PHI during optimization. The benchmark attributes this to four interacting factors: controlled linguistic diversity through explicit variation types, semantic grounding through known canonical values, realistic conversational structure without PHI exposure, and prompt optimization through DSPy/SIMBA instead of manual iterative prompting [2508.15801].

The work also identifies concrete failure modes. DOB extraction is more fragile than ZIP or Name, partly because some synthetic samples use `dd-mm-yyyy` formats that are rarer in US data, creating a domain mismatch at evaluation time. The benchmark also notes model-specific weaknesses, including cases where Gemini 2.5 Pro shows lower synthetic test accuracy. Validation is presented as important for maintaining quality, with invalid samples reported as less than 1%. These observations imply that the quality of the synthetic distribution, not only the optimizer, materially affects downstream transfer [2508.15801].

Several limitations are stated explicitly. Current entity coverage is limited to ZIP, Name, and DOB. Dialogue structure mainly reflects direct answers to questions rather than indirect answers, off-topic utterances, or multi-turn negotiation before the entity is stated. The pipeline depends heavily on LLM quality for both generation and validation, so model biases can propagate into the synthetic corpus. The authors also caution that synthetic data may not fully capture the richness of real human conversations. Proposed future directions include multi-turn dialogues, multilingual synthetic conversational data, explicit incorporation of ASR errors and background noise, human-in-the-loop feedback, and expansion to additional entity types such as medical record numbers, medication names, dosages, and insurance IDs [2508.15801].

Within the broader benchmark landscape, LingVarBench occupies a specific niche. LingBench++ evaluates IOL-style multi-step and cross-cultural linguistic reasoning with structured reasoning traces, typological metadata for over 90 languages, and multi-agent reasoning, whereas LingVarBench targets structured extraction from synthetic healthcare phone transcripts [2507.16809]. GlobalBench is an ever-expanding multilingual and equity-aware benchmarking layer built on ExplainaBoard, with per-speaker utility and Gini-based equity metrics across languages, while LingVarBench is not a language-equity benchmark and does not model typological coverage [2305.14716]. VarBench parameterizes benchmark items through dynamic variable perturbation to mitigate contamination; LingVarBench likewise builds benchmark instances generatively, but its core variables are structured field values and spoken variation types rather than perturbed reasoning templates [2406.17681].

A common source of confusion arises from the name. LingVarBench is not a general benchmark of language variation across languages, dialects, or sociolects in the manner suggested by GlobalBench or by cross-linguistic reasoning benchmarks. Its “variation” is primarily the controlled space of spoken realizations of structured values in synthetic phone-call transcripts. Within that scope, however, it is presented as the first systematic benchmark for structured extraction from synthetic conversational data and as a healthcare-specific methodology for overcoming cost, privacy, and regulatory barriers in large-scale phone call analysis [2508.15801].

Source: https://www.emergentmind.com/topics/lingvarbench