AudioJudge: LAM-Based Audio Evaluation
- AudioJudge is a unified evaluation paradigm that uses large audio models for assessing speech, music, and audio–text alignment via prompt-driven techniques.
- It decomposes evaluation into specialized aspects like lexical content, paralinguistic features, and speech quality using multi-aspect ensemble methods.
- Empirical studies show that AudioJudge achieves high system-level correlations with human judgments, offering robust performance across diverse audio tasks.
AudioJudge denotes a family of automated audio evaluation paradigms in which a large audio model or multimodal model acts as a judge over speech, music, or broader audio–text alignment. In its canonical formulation, AudioJudge is a Large Audio Model (LAM) based framework for speech evaluation that replaces multiple task-specific evaluators with prompted judging over pronunciation, speaking rate, speaker identification, speech quality, and system-level human preference simulation; later work reuses the same judge paradigm for audio-domain similarity in memorization studies and for rubric-based audio instruction-following evaluation (Manakul et al., 17 Jul 2025, Roh et al., 23 Jul 2025, Li et al., 2 Jun 2026).
1. Origin and conceptual scope
The original AudioJudge formulation was introduced as a response to two limitations of conventional speech evaluation: the need to design specialized systems for individual audio characteristics, and the poor correlation of many automatic metrics with human preferences. The central proposal was a unified evaluation approach in which the same LAM serves as a judge across tasks, with the evaluation target determined primarily by the prompt and, where needed, by in-context examples. The framework studies speech-in, text-out judges such as GPT-4o-Audio and Gemini-family models, and also introduces a multi-aspect ensemble that decomposes assessment into lexical content, paralinguistic features, and speech quality (Manakul et al., 17 Jul 2025).
A second, narrower use of the term appears in work on memorization attacks in lyrics-to-song generation. There, AudioJudge is described as a “large audio model based framework that simulates human preferences judgments” and is used as the primary audio-domain similarity metric for comparing generated songs against presumed training songs, particularly under phonetic attacks that preserve acoustic structure while changing lexical meaning. In that setting, AudioJudge is not an embedding distance but a judge model focused on melody and rhythm similarity (Roh et al., 23 Jul 2025).
Subsequent systems extend the broader “audio judge” idea beyond speech quality or memorization. AnyAudio-Judge reframes audio instruction-following evaluation as dynamic rubric-based verification, while SpeechJudge specializes the judge paradigm for speech naturalness using large-scale human preference data and a generative reward model. These systems do not replace the original AudioJudge formulation; rather, they show that the judge role can be specialized by task, supervision, and output format (Li et al., 2 Jun 2026, Zhang et al., 11 Nov 2025).
2. Core architecture and prompting strategies
In the original framework, AudioJudge operates through prompted LAM inference rather than a fixed learned head. System prompts specify the evaluation aspect, user messages attach one or more audio clips, and the output is constrained to JSON with a reasoning field plus a label or match field. The prompting study systematically varies zero-shot versus in-context learning, how multiple audio clips are supplied, and whether transcripts are added. Five audio formatting strategies are reported for pairwise tasks: No Concat, Pair Example Concat, Examples Concat, Test Concat, and Examples+Test Concat. Among these, Test Concat and especially Examples+Test Concat with four shots are reported as the strongest general configurations (Manakul et al., 17 Jul 2025).
The multi-aspect ensemble is the most explicit architectural decomposition. It runs three specialized judges with different prompts over the same input: a Lexical Judge that evaluates content “as if reading transcripts,” a Paralinguistic Judge that focuses on tone, prosody, expressiveness, accent, and related features, and a Speech Quality Judge that focuses on clarity, naturalness, fluency, pronunciation correctness, noise, and artifacts. Final decisions are combined by majority vote. This turns AudioJudge from a single prompted evaluator into a compositional judging system whose sub-judges correspond to distinct speech dimensions (Manakul et al., 17 Jul 2025).
In the music-memorization setting, the prompting scheme is simpler and more scalar. The paper reports the exact system instruction:
COMPARE THE TWO AUDIO FILES BASED ON MELODY AND RHYTHM and how similar they are. Provide a detailed analysis and provide numerical scores (0-1) for each category. Give detailed reasoning for your scores and summarize overall similarity.
Under this protocol, AudioJudge returns Melody and Rhythm scores in , and the authors also visualize aggregate “overall similarity” on a scale in appendix heatmaps. The framework is implemented through GPT-4o-audio-preview and is used at the whole-clip level rather than via explicit segment-level aggregation (Roh et al., 23 Jul 2025).
A later specialization, AnyAudio-Judge, departs from holistic judging by decomposing an instruction into a variable-size set of binary rubric items , estimating a yes/no probability for each item, and averaging those probabilities into a final alignment score . That design makes explicit what the original AudioJudge leaves implicit: evaluation can be factorized into independent, verifiable sub-questions rather than treated as one monolithic judgment (Li et al., 2 Jun 2026).
3. Empirical tasks and reported performance
AudioJudge was evaluated at both example level and system level. Example-level characteristic detection covered pronunciation accuracy, speaking rate detection, speaker identification, and pairwise speech quality assessment on SOMOS, TMHINTQ, and ThaiMOS. With GPT-4o-Audio in a 0-shot, no-concatenation configuration, reported accuracies were 46.0% for pronunciation, 46.9% for speaking rate, 61.5% for speaker identification, 70.5% for SOMOS, 70.5% for TMHINTQ, and 65.5% for ThaiMOS. With the best prompting configuration, reported as 4-shot Examples+Test Concat, these became 66.5%, 55.3%, 70.0%, 71.0%, 74.5%, and 64.0%, respectively. Human accuracies on the characteristic tasks are reported in the approximate range 69.5–83.0%, leaving a notable gap on non-lexical perception, especially speaking rate (Manakul et al., 17 Jul 2025).
At the system level, AudioJudge was used to simulate human preferences over spoken system outputs. On ChatbotArena-Spoken, a lexical benchmark derived from spoken versions of ChatbotArena comparisons, GPT-4o-Audio reached a Spearman correlation of 0.902 in the baseline setting and 0.931 with 4-shot Examples+Test Concat. On SpeakBench, a speech-in speech-out benchmark introduced for pronunciation, speaking style, prosody, delivery, and non-linguistic sound effects, a single GPT-4o-Audio judge reached 0.731 in the baseline setting and 0.775 with in-context learning, while the multi-aspect ensemble improved this to 0.802 in 0-shot and 0.846 with in-context learning. Gemini-2.5-Flash reached 0.912 on SpeakBench with the multi-aspect ensemble in the zero-shot setting (Manakul et al., 17 Jul 2025).
The same study compared AudioJudge against specialized speech-quality predictors. On SOMOS and TMHINTQ, supervised baselines such as MOSANET+ outperformed AudioJudge, but on ThaiMOS AudioJudge achieved 64.0%, exceeding UTMOS, MOSANET+, and SALMONN-FT. The reported interpretation is that AudioJudge trades some in-domain performance for lower annotation overhead and stronger cross-domain behavior (Manakul et al., 17 Jul 2025).
These results delimit the regime in which AudioJudge is strongest. System-level ranking and lexical preference simulation can reach correlations comparable to strong text-judge settings, whereas fine-grained non-lexical detection remains harder. This asymmetry reappears in later multimodal work, where understanding tasks are easier to judge than open-ended generation tasks, especially when the output modality is audio (Pu et al., 21 Mar 2025).
4. AudioJudge in memorization and audio-domain similarity
In “Bob’s Confetti,” AudioJudge is repurposed from speech evaluation to memorization analysis in lyrics-to-song generation. The study introduces Adversarial PhoneTic Prompting (APT), where lyrics are semantically altered but acoustically preserved through homophonic substitutions, such as “mom’s spaghetti” becoming “Bob’s confetti.” AudioJudge functions there as the main audio-domain similarity and evaluation tool for measuring how close the generated song is to a known reference song, alongside CLAP similarity and CoverID (Roh et al., 23 Jul 2025).
The reported AudioJudge scores are high enough to support a memorization interpretation. For Kendrick Lamar’s “DNA,” phoneme-modified SUNO generations produced Melody and Rhythm in two independent generations. For “Lose Yourself,” a phoneme-variant prompt with “Bob’s confetti” yielded Melody and Rhythm with an “intense rap” genre prompt, and Melody , Rhythm 0 without genre conditioning. For “Jingle Bell Rock,” multiple phoneme-substituted variants consistently produced Melody 1 with Rhythm in the range 2. The authors interpret values around 3 as near-identical melody or rhythm, 4 as strong similarity, and values around 5 or below as low similarity (Roh et al., 23 Jul 2025).
The same pattern extends across models, languages, and genres. In SUNO, an APT parody of “APT” by ROSÉ and Bruno Mars reached Melody 6 and Rhythm 7. In YuE, exact-lyrics reuse and cross-lingual examples produced similarly high scores: “Let It Be” reached Melody 8, Rhythm 9, and Teresa Teng’s “月亮代表我的心” reached Melody 0, Rhythm 1. Exact-lyrics YuE evaluations for Mandarin and Cantonese songs, such as “红日,” “光辉岁月,” and “海阔天空,” produced melody scores in the 2 range and rhythm scores in the 3 range. Appendix heatmaps aggregate these pairwise judgments into 4 overall similarity values, with matched original–variant pairs often at 80–96 and unrelated pairs frequently in the 12–25 range (Roh et al., 23 Jul 2025).
Within that paper, AudioJudge is more than an auxiliary metric. It is the main evidence that phonetic structure alone can trigger musically near-duplicate outputs. The authors use those audio findings to motivate a broader multimodal claim, “phonetic-to-visual regurgitation,” in which phoneme-altered lyrics also trigger reconstruction of visual elements in text-to-video generation. AudioJudge itself remains audio-domain, but it anchors the paper’s argument that memorization can be sub-lexical and rhythm-driven rather than limited to exact textual matches (Roh et al., 23 Jul 2025).
5. Specialized descendants and neighboring judge systems
The development of audio judges after AudioJudge has proceeded in two main directions: specialization by task and formalization by supervision.
Before the table below, one distinction is central. AudioJudge in the original sense is a prompting framework over existing LAMs; AnyAudio-Judge is a trained evaluator and benchmark for instruction following; SpeechJudge is a naturalness-focused suite centered on human preference data and a generative reward model. All three retain the idea that a model can act as a human-aligned evaluator, but they differ in decomposition granularity, supervision, and target construct.
| System | Target task | Notable report |
|---|---|---|
| AudioJudge | Speech evaluation and system ranking | Up to 0.91 Spearman correlation with human preferences; multi-aspect ensemble; audio concatenation + in-context learning (Manakul et al., 17 Jul 2025) |
| AnyAudio-Judge | Audio instruction-following alignment | 7,920-sample bilingual benchmark; 105K-sample corpus; 85.26 ACC on Chinese and 84.45 ACC on English (Li et al., 2 Jun 2026) |
| SpeechJudge | Speech naturalness judgment | 99K speech pairs; 77.2% accuracy, 79.4% with inference-time scaling @10 (Zhang et al., 11 Nov 2025) |
AnyAudio-Judge introduces a dynamic rubric-based evaluation paradigm. It decomposes an instruction into atomic yes/no questions, predicts logits for “yes” and “no” for each item, converts them into soft satisfaction probabilities, and averages them into an alignment score 5. The accompanying AnyAudio-Judge Bench comprises 7,920 balanced positive and negative samples across speech, sound, music, and mixed domains, with bilingual English and Chinese versions and deliberately constructed hard negatives. The trained evaluator is built on Qwen3-Omni-30B-A3B-Captioner and is trained using Supervised Fine-Tuning and Group Relative Policy Optimization. Relative to strong general-purpose baselines, it reports 85.26 average ACC on the Chinese subset and 84.45 on the English subset, and on PAM it reports LCC 6, SRCC 7, and KTAU 8 (Li et al., 2 Jun 2026).
SpeechJudge narrows the target construct to naturalness in zero-shot TTS. It introduces SpeechJudge-Data with 99K speech pairs, SpeechJudge-Eval with 1,000 high-agreement evaluation pairs, and SpeechJudge-GRM, a generative reward model based on Qwen2.5-Omni-7B. On SpeechJudge-Eval, reported accuracies include 57.9% for WER, 53.7% for UTMOS, 67.4% for GPT-4o Audio, 69.1% for Gemini-2.5-Flash, 72.7% for a Bradley–Terry reward model baseline, 75.3% for SpeechJudge-GRM after SFT, and 77.2% after SFT+RL, rising to 79.4% with Voting@10. The model is then used as a reward function for post-training a TTS generator, where the online GRM-guided setting yields the largest reported naturalness CMOS gain, 9 (Zhang et al., 11 Nov 2025).
A broader multimodal context is provided by JudgeAnything, which evaluates MLLM-as-a-judge performance across text, image, video, audio, and audiovisual tasks. Its reported aggregate results show that judges are substantially more reliable on multimodal understanding than on multimodal generation, and that output modality matters more than input modality, with audio outputs among the harder cases. This situates AudioJudge within a general finding: automated judging is currently easier when the judge can reduce the problem to understanding rather than scoring open-ended generated audio (Pu et al., 21 Mar 2025).
6. Biases, calibration, and methodological implications
The original AudioJudge study reports two favorable and two unfavorable properties. On the favorable side, the judges maintain strong performance under acoustic noise, and with the right prompting they can achieve high system-level correlation with human preferences. On the unfavorable side, robustness analysis reveals significant verbosity bias and positional bias. Verbosity bias appears when judges prefer longer responses in ambiguous lexical cases; positional bias appears when decisions change systematically with order reversal or favor the first or second slot independently of content. These effects are modest enough not to destroy system-level ranking on the reported benchmarks, but they remain central caveats for deployment (Manakul et al., 17 Jul 2025).
Later psychometric work argues that judge models should be treated as measurement instruments rather than summarized only by scalar agreement. The proposed Judge Datasheet protocol introduces dark current under true-vacuum inputs, stable cross-sensitivity to same-quality surface variation, positional false preference, target sensitivity on a controlled quality ladder, and the criterion or operating point induced by tie instructions. Its main conclusion is that prompting moves the criterion, not the resolution: stricter tie instructions can eliminate false preferences on same-quality pairs, but often by converting marginal true differences into ties rather than by improving the underlying discriminative capacity (Usami et al., 14 Jun 2026).
This framework is directly relevant to AudioJudge-style systems. A plausible implication is that an audio judge should be characterized not only by benchmark accuracy or rank correlation, but also by behavior on identical clips, silence or “true vacuum” inputs, order-swapped pairings, and controlled quality ladders. That implication is reinforced by JudgeAnything, where checklist prompting sometimes improves strong judges but can also induce hallucination or tie overuse, especially in audio and video generation tasks (Usami et al., 14 Jun 2026, Pu et al., 21 Mar 2025).
Taken together, the literature places AudioJudge at the intersection of evaluation, alignment, and safety. In speech evaluation it provides a unified LAM-based alternative to collections of specialized models; in music-generation security it serves as a human-like similarity metric for memorization analysis; in later systems it becomes a rubric-based evaluator or a trained reward model. The unifying theme is not a single architecture, but the substitution of human-like auditory judgment with a structured model-based instrument whose capabilities, biases, and calibration must themselves be measured (Manakul et al., 17 Jul 2025, Roh et al., 23 Jul 2025, Li et al., 2 Jun 2026).