Papers
Topics
Authors
Recent
Search
2000 character limit reached

LingVarBench: Synthetic NER Benchmark for Healthcare Calls

Updated 8 July 2026
  • LingVarBench is a synthetic benchmark and pipeline that generates and validates structured conversational data for healthcare NER tasks.
  • It employs a three-stage generative process—value generation, transcript creation, and validation—to capture real phone-call disfluencies and variations.
  • Optimized prompts via DSPy and SIMBA enhance extraction accuracy while bypassing the high costs and privacy issues of real transcript labeling.

LingVarBench is a synthetic benchmark and pipeline for training and evaluating LLM–based named entity recognition and structured extraction systems on phone-call style conversational speech, particularly in healthcare settings such as patient–provider calls involving ZIP codes, dates of birth, and names. It addresses the problem that labeling real phone transcripts is expensive, privacy-sensitive, and operationally constrained, while conventional NER benchmarks do not reflect the disfluencies, interruptions, informal syntax, and formatting variability of conversational speech. In this formulation, LingVarBench is both a benchmark and a method: it generates structured synthetic utterances, validates recoverability of the underlying entities, and uses the resulting data to optimize extraction prompts that transfer to real customer transcripts (Mohammadi et al., 13 Aug 2025).

1. Problem formulation and scope

LingVarBench is motivated by the need to extract structured information from spontaneous speech in healthcare voice AI systems. The target data regime is phone call transcripts containing disfluencies, self-corrections, hesitations and pauses, filler words, variable formatting, background noise and ASR errors in real data, and domain-specific healthcare terminology and emotional content. The benchmark focuses on fields that are operationally important for authentication and downstream workflow, including ZIP code, Date of Birth, and Name (Mohammadi et al., 13 Aug 2025).

The central obstacle is the scarcity of large, high-quality, HIPAA-compliant labeled datasets for conversational healthcare speech. The reported constraints are concrete: phone call transcript labeling is “approximately 2 USD per minute,” and manual annotation requires “3 hours of expert time per hour of audio.” The same setting also introduces privacy and governance burdens, including PHI handling, consent requirements, BAAs, and privacy review. LingVarBench is designed as a way to bypass these barriers by replacing direct annotation of real transcripts with validated synthetic conversational data and automated prompt optimization (Mohammadi et al., 13 Aug 2025).

This design distinguishes LingVarBench from standard written-text NER benchmarks such as CoNLL-2003, OntoNotes, and clinical NER over written notes. Those resources do not capture phone-call style conversational phenomena and are not constructed under the same regulatory constraints. A further distinction is that LingVarBench is domain-specific rather than multilingual in the sense used by broad equity-oriented benchmarks such as GlobalBench, which tracks 966 datasets in 190 languages and evaluates utility and equity across languages (2305.14716). LingVarBench instead concentrates on structured extraction under conversational variation in a healthcare workflow.

2. Generative architecture and validation loop

LingVarBench is organized as a three-stage generative pipeline followed by prompt optimization. The stages are a Value Generator VV, a Transcript Generator TT, and a Validator C\mathcal{C}. The Value Generator produces plausible canonical structured values for a field description; the Transcript Generator transforms those values into natural conversational utterances under specified variation types; and the Validator retains only utterances from which the target value can be recovered. The validated corpus is then used inside DSPy with SIMBA to optimize extraction prompts (Mohammadi et al., 13 Aug 2025).

Formally, the paper defines entity types e∈Ee \in \mathcal{E}, field descriptions d∈Dd \in \mathcal{D}, variation types v∈Vv \in \mathcal{V}, and utterances u∈Uu \in \mathcal{U}, with component mappings

V:D→V,T:(V,V)→U,C:U→(U,e).V: \mathcal{D} \rightarrow \mathbb{V}, \qquad T: (\mathbb{V}, \mathcal{V}) \rightarrow \mathcal{U}, \qquad \mathcal{C}: \mathcal{U} \rightarrow (\mathcal{U}, e).

The composite generator is

G(d,v)=C(T(V(d), v)),G:(D,V)→(U,e).G(d, v) = \mathcal{C}\left(T\left(V(d),\, v\right)\right), \quad G : (\mathcal{D}, \mathcal{V}) \rightarrow (\mathcal{U}, e).

This means that, for each field description and variation type, LingVarBench produces a validated pair (u,e)(u,e): a transcript containing a recoverable entity instance (Mohammadi et al., 13 Aug 2025).

A concise view of the workflow is as follows:

Stage Function Output
Value Generator TT0 Samples plausible canonical values from a field description Structured values such as “91101” or “12-02-1947”
Transcript Generator TT1 Verbalizes a canonical value under specified variation types Spoken-style transcripts with variation tags
Validator TT2 Tests whether the value is recoverable from the transcript Filtered transcript–entity pairs
DSPy + SIMBA Optimizes extraction prompts on validated synthetic data Entity-specific prompts for real and synthetic evaluation

The validator is a critical quality-control mechanism. Generated utterances are not accepted by default; they are checked by a separate LLM acting as a binary extractor or verifier. If the transcript contains the value, even with corrections, validation returns true; if it is vague or does not contain the value, it returns false. For dates, the validation logic explicitly accepts both continuous digits and spoken sequences. Invalid transcripts are discarded, and generation may be recursively re-invoked to fill the quota for a value–variation pairing. The benchmark is therefore not merely synthetic generation but “generate, verify, filter” (Mohammadi et al., 13 Aug 2025).

3. Linguistic variation taxonomy and dataset construction

A defining feature of LingVarBench is its explicit taxonomy of linguistic variation types. Each field description comprises a field name, a data type, and a natural-language description specifying format or constraints. The transcript generation stage then conditions on one or more variation types that model how structured values are realized in speech (Mohammadi et al., 13 Aug 2025).

The variation taxonomy has two layers. First, there are general variations that can apply to any entity, including filler_words, hesitation, correction or self-correction, repetition, pause, formal, casual, polite, confident, uncertain, rushed, careful, confirmation, and direct_and_simple. Second, there are entity-specific variations. For ZIP code, these include digit_by_digit, grouped_two, grouped_three, hundred, mixed_grouping, and reversed, along with combinations involving hesitations, fillers, and corrections. For DOB, the taxonomy includes date_as_4_digits, date_as_5_digits, date_as_6_digits, date_as_8_digits, spoken_date_4_digits through spoken_date_8_digits, spoken_month_day_year, mixed_spoken_and_digits, filler_or_correction, and casual_or_polite_digits. For Name, the benchmark includes name_with_last, name_with_prefix, name_reverse_order, name_with_title, name_with_middle, name_with_suffix, name_with_initials, name_with_correction, name_partial_spelling, name_with_apostrophe, name_hyphenated, and nickname (Mohammadi et al., 13 Aug 2025).

The transcript generator pairs canonical values with these variation types, generates utterances in JSON form, and tracks tag frequencies. To avoid concentration on easy or highly frequent realizations, the system re-prompts for underrepresented value–variation pairs until a balanced distribution is reached. The benchmark description states that this yields “uniform coverage over linguistic variation types while preserving diversity at the utterance level” (Mohammadi et al., 13 Aug 2025).

After generation and validation, the corpus is split into canonical train, validation, and test sets at 70% / 15% / 15%. The reported summary statistics show that the benchmark is substantial even in its initial scope:

Entity GPT-4 synthetic set Gemini 2.0 Flash synthetic set
ZIP code 5,635 samples; 33 variation tags; 10 unique canonical values; avg length ~41 tokens 6,467 samples; 31 tags; 12 values
Name 2,550 samples; 46 tags; 12 values 6,631 samples; 46 tags; 12 values
DOB 1,055 samples; 36 tags; 7 values 682 samples; 33 tags; 5 values

The corpus structure makes LingVarBench a benchmark rather than only a generation procedure: each sample has a transcript, a ground-truth canonical value, and associated variation tags. The benchmark is therefore suitable for both supervised prompt optimization and controlled robustness analysis across conversational phenomena (Mohammadi et al., 13 Aug 2025).

4. DSPy, SIMBA, and automated prompt optimization

The synthetic corpus is used as training signal for automated prompt optimization rather than as an end in itself. LingVarBench integrates DSPy, described as a framework for programming LLM pipelines declaratively and automatically optimizing prompts and hyperparameters, with SIMBA, the Stochastic Introspective Mini-Batch Ascent optimizer. The explicit objective is to eliminate manual prompt engineering while learning prompts that are robust to conversational variation (Mohammadi et al., 13 Aug 2025).

For each entity type, the procedure is: generate synthetic labeled transcripts; define a base extraction module in DSPy with general and field-specific instructions; optimize the prompt on synthetic training and validation data; and then evaluate the resulting prompt on both synthetic held-out data and real phone transcripts. The benchmark description gives representative field-specific instructions such as “Extract exactly 5 numeric digits,” “Do not interpret ZIP codes as dates,” and “Return the raw 5-digit number only” (Mohammadi et al., 13 Aug 2025).

The extractor is treated as a prompt-parameterized function TT3. The optimization target is stated either as exact-match accuracy,

TT4

or equivalently as F1,

TT5

SIMBA then performs mini-batch, introspective updates over candidate prompts, evaluating rephrasings and constraints on synthetic data and ascending toward prompts with better extraction performance (Mohammadi et al., 13 Aug 2025).

This optimization strategy is central to the benchmark’s claim of practicality. The prompt search uses only synthetic validated transcripts and therefore avoids using PHI-containing real transcripts during optimization. In that sense, LingVarBench is not merely a synthetic dataset for model pretraining; it is a closed-loop system for generating task-specific training data, validating semantic faithfulness, and using the result to optimize LLM extraction behavior under privacy constraints (Mohammadi et al., 13 Aug 2025).

5. Experimental setup and empirical results

Generation, validation, and extraction are instantiated with GPT-4, Gemini 2.0 Flash, and Gemini 2.5 Pro, using standardized prompt formats and deterministic decoding with temperature TT6 and top-p TT7. Similarity analysis uses text-embedding-3-large and gemini-embedding-001. Evaluation is conducted on both the synthetic LingVarBench splits and a proprietary real dataset from Infinitus Systems Inc. consisting of real patient-facing phone calls reviewed in a human-in-the-loop process for ZIP code, Name, and DOB ground truth (Mohammadi et al., 13 Aug 2025).

The baseline comparison includes three prompt conditions: a zero-shot prompt, a human prompt iteratively refined with real data, and a LingVarBench-optimized prompt obtained through DSPy/SIMBA on synthetic data alone. On synthetic data, the optimized prompts achieve high accuracy: ZIP code reaches approximately 97% for GPT-4 and approximately 89% for Gemini 2.0 Flash; Name reaches approximately 96–97% for GPT-4 and approximately 79% for Gemini 2.0 Flash; DOB reaches approximately 98% for GPT-4, while Gemini 2.5 Pro reaches 100% on validation and approximately 98% on test (Mohammadi et al., 13 Aug 2025).

The principal reported results concern transfer to real transcripts. For ZIP code, zero-shot performance is 88.62% accuracy for GPT-4 and 89.37% for Gemini 2.5 Pro; the best human prompt reaches 89.76% for Gemini 2.0 Flash; and LingVarBench-optimized prompting reaches 95.07% accuracy and F1 94.61 for Gemini 2.5 Pro, with GPT-4 at 89.79% accuracy and F1 94.66. For Name, zero-shot ranges from 47.87% accuracy for Gemini 2.0 Flash to 78.19% for GPT-4, whereas LingVarBench-optimized prompts reach 90.18% accuracy and F1 94.84 for both GPT-4 and Gemini 2.0 Flash. For DOB, zero-shot is 74.52% for GPT-4 and 77.04% for Gemini 2.5 Pro; human prompting reaches 82.62% for Gemini 2.0 Flash; and LingVarBench-optimized prompting reaches 80.80% accuracy and F1 89.38 for Gemini 2.5 Pro, and 80.66% accuracy and F1 89.30 for Gemini 2.0 Flash (Mohammadi et al., 13 Aug 2025).

The paper summarizes these improvements as up to 95 percent accuracy for numeric fields versus 88–89 percent zero-shot, around 90 percent for names versus 47–79 percent zero-shot, and over 80 percent for dates versus 72–77 percent zero-shot. The resulting interpretation is narrow but strong: optimized synthetic-data prompts consistently outperform zero-shot prompts, match or surpass human-tuned prompts for ZIP and Name, and improve over zero-shot for DOB, although human-tuned prompts may remain slightly stronger in some DOB settings (Mohammadi et al., 13 Aug 2025).

A further evaluation axis is synthetic-to-real similarity. Using cosine similarity over embeddings, when variation types match, text-embedding-3-large yields 0.81 ± 0.13 for ZIP, 0.81 ± 0.15 for Name, and 0.62 ± 0.14 for DOB; gemini-embedding-001 yields 0.91 ± 0.07 for ZIP, 0.92 ± 0.08 for Name, and 0.83 ± 0.08 for DOB. Similarity drops when variation tags diverge, which the paper uses to argue that explicit variation control makes synthetic utterances closer to real conversational patterns (Mohammadi et al., 13 Aug 2025).

6. Interpretation, limitations, and relation to adjacent benchmarks

The main analytical claim of LingVarBench is that controlled linguistic diversity plus automated prompt optimization can produce robust extractors for real conversational speech without using real PHI during optimization. The benchmark attributes this to four interacting factors: controlled linguistic diversity through explicit variation types, semantic grounding through known canonical values, realistic conversational structure without PHI exposure, and prompt optimization through DSPy/SIMBA instead of manual iterative prompting (Mohammadi et al., 13 Aug 2025).

The work also identifies concrete failure modes. DOB extraction is more fragile than ZIP or Name, partly because some synthetic samples use dd-mm-yyyy formats that are rarer in US data, creating a domain mismatch at evaluation time. The benchmark also notes model-specific weaknesses, including cases where Gemini 2.5 Pro shows lower synthetic test accuracy. Validation is presented as important for maintaining quality, with invalid samples reported as less than 1%. These observations imply that the quality of the synthetic distribution, not only the optimizer, materially affects downstream transfer (Mohammadi et al., 13 Aug 2025).

Several limitations are stated explicitly. Current entity coverage is limited to ZIP, Name, and DOB. Dialogue structure mainly reflects direct answers to questions rather than indirect answers, off-topic utterances, or multi-turn negotiation before the entity is stated. The pipeline depends heavily on LLM quality for both generation and validation, so model biases can propagate into the synthetic corpus. The authors also caution that synthetic data may not fully capture the richness of real human conversations. Proposed future directions include multi-turn dialogues, multilingual synthetic conversational data, explicit incorporation of ASR errors and background noise, human-in-the-loop feedback, and expansion to additional entity types such as medical record numbers, medication names, dosages, and insurance IDs (Mohammadi et al., 13 Aug 2025).

Within the broader benchmark landscape, LingVarBench occupies a specific niche. LingBench++ evaluates IOL-style multi-step and cross-cultural linguistic reasoning with structured reasoning traces, typological metadata for over 90 languages, and multi-agent reasoning, whereas LingVarBench targets structured extraction from synthetic healthcare phone transcripts (Lian et al., 22 Jul 2025). GlobalBench is an ever-expanding multilingual and equity-aware benchmarking layer built on ExplainaBoard, with per-speaker utility and Gini-based equity metrics across languages, while LingVarBench is not a language-equity benchmark and does not model typological coverage (2305.14716). VarBench parameterizes benchmark items through dynamic variable perturbation to mitigate contamination; LingVarBench likewise builds benchmark instances generatively, but its core variables are structured field values and spoken variation types rather than perturbed reasoning templates (Qian et al., 2024).

A common source of confusion arises from the name. LingVarBench is not a general benchmark of language variation across languages, dialects, or sociolects in the manner suggested by GlobalBench or by cross-linguistic reasoning benchmarks. Its “variation” is primarily the controlled space of spoken realizations of structured values in synthetic phone-call transcripts. Within that scope, however, it is presented as the first systematic benchmark for structured extraction from synthetic conversational data and as a healthcare-specific methodology for overcoming cost, privacy, and regulatory barriers in large-scale phone call analysis (Mohammadi et al., 13 Aug 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to LingVarBench.