Papers
Topics
Authors
Recent
Search
2000 character limit reached

WavBench: Benchmark for End-to-End Dialogue

Updated 14 July 2026
  • WavBench is a large-scale benchmark that evaluates spoken dialogue models on reasoning, colloquialism, and paralinguistic fidelity.
  • It consists of 17,577 audio items over 76.5 hours divided into Pro, Basic, and Acoustic Interaction subsets to reveal limitations of traditional text-based evaluations.
  • The benchmark employs a multi-stage construction pipeline and distinct evaluation protocols to measure content accuracy, natural speech delivery, and prosodic alignment.

WavBench is a large-scale benchmark for end-to-end spoken dialogue models that is explicitly designed to evaluate three capabilities in combination: high-difficulty reasoning, spoken colloquialism, and paralinguistic fidelity within realistic conversational settings (Li et al., 12 Feb 2026). It was introduced in response to the transition from cascaded speech pipelines such as ASR→LLM→TTS toward reasoning-enhanced end-to-end systems, including Moshi, GLM-4-Voice, and Kimi-Audio, where speech is no longer treated merely as text with a vocal rendering but as the medium in which models are expected to “think and talk” (Li et al., 12 Feb 2026). The benchmark comprises 17,577 audio items spanning 76.5 hours and is organized into three subsets—Pro, Basic, and Acoustic Interaction—intended to expose limitations that text-centric or narrowly acoustic evaluations do not adequately capture (Li et al., 12 Feb 2026).

1. Scope and motivation

WavBench is motivated by two observations. First, contemporary spoken dialogue models increasingly integrate advanced reasoning capabilities directly into end-to-end audio systems. Second, existing evaluations predominantly follow text-generation standards and therefore overlook audio-centric properties such as paralinguistics and colloquialisms, as well as the cognitive depth expected of modern spoken agents (Li et al., 12 Feb 2026).

The benchmark is framed as a response to specific deficiencies in prior evaluation practice. According to its formulation, existing benchmarks for speech comprehension, task-oriented dialogue, acoustic attributes, or text-adapted reasoning either neglect fluency and rapport, ignore implicit prosody, or treat audio simply as a vessel for text (Li et al., 12 Feb 2026). WavBench therefore measures three orthogonal dimensions within real-world scenarios: high-difficulty reasoning with spoken colloquialism, everyday conversational vivacity, and comprehensive paralinguistic fidelity (Li et al., 12 Feb 2026).

This design places particular emphasis on spoken interaction as a distinct target of evaluation rather than as a speech-transcribed proxy for text generation. In the benchmark’s own framing, the intended challenges are to simplify multi-step proofs or algorithms into listener-friendly speech; maintain lexical appropriateness, linguistic naturalness, and interactive rapport in routine dialogue; and perceive and generate fine-grained acoustic attributes under both explicit instructions and implicit multi-turn conversational conditions (Li et al., 12 Feb 2026). This suggests a deliberate shift away from benchmarks dominated by written-form correctness alone.

2. Benchmark architecture

WavBench is divided into three complementary subsets. Together they cover difficult reasoning, low-to-medium complexity conversational fluency, and explicit as well as implicit paralinguistic behavior (Li et al., 12 Feb 2026).

Subset Size Primary focus
Pro 3,176 items High-cognitive-load “Stress Tests” in seven domains
Basic 4,486 items “Everyday Fluency” for low-to-medium complexity dialogue
Acoustic Interaction 9,915 items Explicit and implicit paralinguistic evaluation

The Pro subset is described as stressing high-cognitive-load scenarios in seven domains: Code, Creative Writing, Instruction Following, Logic, Math, Common QA, and Safety (Li et al., 12 Feb 2026). Its source data derive from 15 public text corpora, with the reported composition including Arena-Hard (35%), MMLU (25%), BBEH (35%), GPQA (6%), Math benchmarks (5%), COLLIE (4%), and Code tasks (2%), followed by stratification with GPT-4.1 into high-difficulty and distractor examples (Li et al., 12 Feb 2026). Each item contains a synthetic user question and a reference colloquial response, and the final audio is recorded in 1,088 diverse speaker voices after spoken-query rewriting, response colloquialization, human verification of 11K samples, and IndexTTS2 speech synthesis with WER<5%WER < 5\% (Li et al., 12 Feb 2026).

The Basic subset also spans the same seven cognitive domains but is aimed at low-to-medium complexity tasks where the purpose is not deep reasoning stress but assessment of whether a model sounds “alive” and “friendly” in routine exchanges (Li et al., 12 Feb 2026). Its data are drawn from OpenBookQA (25%), WildSpeech (23%), AlpacaEval (18%), AlignBench (13%), MMLU-ProX (12%), plus Math (7%) and minor contributions from BBEH and AutoLogi (1% each) (Li et al., 12 Feb 2026). Construction is reported as following the same five-stage pipeline and producing TTS-synthesized audio with controlled acoustic diversity (Li et al., 12 Feb 2026).

The Acoustic Interaction subset evaluates 10 paralinguistic attributes: four speaker-info dimensions—age, gender, accent, language—four basic acoustics—pitch, speed, volume, emotion—and two background categories—audio events and music (Li et al., 12 Feb 2026). It is further divided into Explicit Understanding and Generation, which are single-turn and directly cued, and Implicit Multi-Turn Dialogue, which uses no direct cues and requires models to infer and match the user’s prosodic context across turns (Li et al., 12 Feb 2026).

3. Construction methodology and data design

The benchmark’s construction pipeline is central to its intended realism. For the Pro subset, the reported process consists of four stages: spoken-query rewriting, response colloquialization, human verification of 11K samples, and IndexTTS2 speech synthesis together with a WER<5%WER < 5\% condition (Li et al., 12 Feb 2026). For the Basic subset, construction is described as following the same five-stage pipeline, again with TTS-synthesized audio and controlled acoustic diversity (Li et al., 12 Feb 2026). The data description does not enumerate each stage separately for Basic beyond this characterization.

The colloquialization objective is not defined as mere stylistic relaxation. Rather, WavBench formalizes spoken quality around “listenability,” emphasizing natural vocabulary, linguistic fluency, and interactive rapport instead of rigid written accuracy (Li et al., 12 Feb 2026). In practical terms, the benchmark specifies lexical appropriateness through common vocabulary and discourse markers, linguistic naturalness through short flexible utterances and natural omissions, and interactive rapport through rhetorical questions and confirmations (Li et al., 12 Feb 2026). This operationalization is particularly important because it distinguishes spoken adequacy from the conventions of written-form answer evaluation.

The Acoustic Interaction subset is designed to move beyond explicit command following toward context-sensitive prosodic adaptation. In its explicit setting, prompts may directly specify a target style, such as “Please adopt an angry tone,” or ask for perceptual recognition, such as “Can you perceive my accent?” (Li et al., 12 Feb 2026). In its implicit setting, the dialogue has a 1:3 single-turn:multi-turn ratio, consists of four-turn dialogues, and contains utterances of 4–25 seconds, with no direct cue regarding the required style (Li et al., 12 Feb 2026). Models must therefore infer the prosodic context from preceding interaction rather than from an overt instruction.

The attribute distribution is also reported in concrete terms. Emotion constitutes the largest slice at approximately 25% of attributes; accents span six varieties—Ind., Can., Brit., Sing., US, Aus.; age covers four brackets from child to elderly; and background events range from wind noise to piano music (Li et al., 12 Feb 2026). These choices indicate a benchmark design that targets both socially salient and acoustically fine-grained distinctions.

4. Evaluation protocol

WavBench uses distinct evaluation procedures for colloquial expression, explicit acoustic tasks, and implicit interaction (Li et al., 12 Feb 2026). The protocols are aligned to the three-part benchmark structure and reflect the authors’ attempt to evaluate spoken dialogue on dimensions not reducible to exact-match or BLEU-like scoring (Li et al., 12 Feb 2026).

For the Pro and Basic subsets, colloquial expression is scored hierarchically by Gemini 3 Pro Preview. Each response receives a score of 1, 3, or 5, where 1 denotes failure, 3 denotes a correct but stiff response, and 5 denotes a correct response that also satisfies all four spoken criteria (Li et al., 12 Feb 2026). The average score per domain is defined as

AvgScore=1Ni=1Nsi,si{1,3,5}.\text{AvgScore} = \frac{1}{N}\sum_{i=1}^{N} s_i,\qquad s_i \in \{1,3,5\}.

For acoustic explicit understanding and generation, the benchmark uses simple accuracy:

Accuracy=1Ni=1N1[y^i=yi].\text{Accuracy} = \frac{1}{N}\sum_{i=1}^{N} 1[\hat{y}_i = y_i].

This formulation is used for the single-turn tasks in which the model either identifies or produces specified paralinguistic attributes (Li et al., 12 Feb 2026).

For implicit interaction, the evaluation is joint over style and content. Paralinguistic style is scored from 1 to 10 via Gemini prompts, and transcription quality is also scored from 1 to 10 (Li et al., 12 Feb 2026). The final implicit score per sample is the mean of style and content ratings (Li et al., 12 Feb 2026). This implies that strong semantic continuation alone is insufficient for high performance if prosodic alignment is absent.

A plausible implication of this metric design is that WavBench does not treat speech output quality as a single scalar property. Instead, it decomposes conversational competence into correctness, spoken-form appropriateness, explicit acoustic controllability, and multi-turn paralinguistic coherence.

5. Experimental setup and benchmarked systems

The reported evaluation covers five state-of-the-art spoken dialogue models, all tested with no additional fine-tuning and under default inference settings (Li et al., 12 Feb 2026). The benchmarked systems are Qwen3-Omni-30B-A3B-Instruct, Kimi-Audio-7B-Instruct, MiMo-Audio-7B-Instruct, Step-Audio-2-mini, and GPT-4o Audio (Li et al., 12 Feb 2026).

Several model-specific descriptors are included in the benchmark report. Qwen3-Omni-30B-A3B-Instruct is characterized as “MoE thinking-talking” with 234 ms latency; Kimi-Audio-7B-Instruct as using a flow-matching detokenizer with 13 M h pretrain; MiMo-Audio-7B-Instruct as using dual-rate tokenization with 100 M h pretrain; Step-Audio-2-mini as using a latent audio encoder with RL for paralinguistics; and GPT-4o Audio as a proprietary OpenAI API system (Li et al., 12 Feb 2026).

The input and auxiliary tooling are also specified. Audio prompts were streamed from IndexTTS2, and Whisper-Large-V3 was used for transcripts in implicit tasks (Li et al., 12 Feb 2026). No ablations were reported (Li et al., 12 Feb 2026). This absence is notable because the benchmark is explicitly multifactorial; however, the published setup focuses on comparative system-level outcomes rather than isolating the contribution of individual design variables.

6. Empirical findings

The Pro subset results indicate that GPT-4o Audio leads with an average score of 58.23, while open-source models remain around 30–40 points (Li et al., 12 Feb 2026). The most difficult areas are reported to be Logic and Math, with Logic dropping to 22.4 for Step-Audio-2 and Math falling in the 25.7–38.6 range across open-source models (Li et al., 12 Feb 2026). These results are presented as evidence that high-cognitive-load spoken reasoning remains a major bottleneck.

On the Basic subset, GPT-4o Audio again ranks highest with 68.80, while Qwen3-Omni scores 55.80 and the remaining systems range from 48.5 to 49.6 (Li et al., 12 Feb 2026). Structured tasks, particularly Math and Instruction, are reported as the hardest even under the lower-complexity framing of the Basic split (Li et al., 12 Feb 2026). This suggests that colloquial delivery alone does not remove the challenge of maintaining procedural or structured correctness in speech.

In acoustic explicit understanding, Step-Audio-2-mini achieves the best average at 57.36%, with especially strong performance in music at 77.8% and language at 96.5% (Li et al., 12 Feb 2026). At the same time, all models score below 35% on pitch, volume, and accent (Li et al., 12 Feb 2026). The explicit understanding results therefore reveal a sharp unevenness across attribute types: some properties are comparatively tractable, while others remain poorly captured.

In acoustic explicit generation, GPT-4o Audio reaches 79.23% average accuracy, with above 95% on gender, emotion, and age, but below 50% on background audio (Li et al., 12 Feb 2026). The asymmetry between speaker-related attributes and background conditions indicates that controllable generation of environmental context is substantially weaker than generation of some salient vocal properties.

The implicit multi-turn setting yields a distinct pattern. Semantic scores rise to 4.4–4.9 out of 10, but style scores collapse to 1.04–1.25 out of 10 (Li et al., 12 Feb 2026). Qwen3-Omni and GPT-4o share the highest combined implicit score at 2.78 (Li et al., 12 Feb 2026). The paper describes a “Cognitive-Acoustic Alignment” gap: models either do logic well or speech style well, but rarely both (Li et al., 12 Feb 2026). This is one of the benchmark’s central empirical conclusions.

7. Significance, limitations, and future directions

WavBench positions itself against a text-centric evaluation regime, including metrics such as BLEU and exact match, by asserting that spoken dialogue models should be evaluated on their ability to reason in colloquial spoken form and to manage paralinguistic context (Li et al., 12 Feb 2026). In that sense, its main significance lies not only in the introduction of new test items but in the redefinition of what constitutes competent spoken interaction for end-to-end systems.

The benchmark’s findings motivate three future directions identified by its authors. The first is reasoning verbalization: training with rich, spoken-style reasoning transcripts so that models learn how to pace, chunk, and colloquialize multi-step solutions (Li et al., 12 Feb 2026). The second is paralinguistic diversity: augmenting corpora with fine-grained prosody labels, including pitch contours, volume dynamics, and environmental sound mixing, to reduce the gap in explicit acoustic tasks (Li et al., 12 Feb 2026). The third is multi-turn coherence: designing dialogue data in which prosodic context evolves naturally, such as a shift from neutral to surprised to sad, so that models learn to track and maintain emotional context across turns (Li et al., 12 Feb 2026).

No ablation studies are reported, and the benchmark is therefore primarily diagnostic rather than mechanistically explanatory (Li et al., 12 Feb 2026). A plausible implication is that WavBench is better suited to identifying system-level failure modes than to determining which architectural or data interventions are most responsible for performance gains. Even so, the reported results establish a clear separation between semantic competence and paralinguistic competence, especially in multi-turn interaction, where style tracking is substantially weaker than content continuation (Li et al., 12 Feb 2026).

The benchmark dataset, transcripts, human-verification scripts, and evaluation prompts are open source through the project website listed in the paper (Li et al., 12 Feb 2026). Within the broader landscape of spoken dialogue research, WavBench functions both as a diagnostic benchmark and as a structured agenda for the development of end-to-end agents that can reason, speak colloquially, and sustain paralinguistic fidelity in realistic conversation (Li et al., 12 Feb 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to WavBench.