Prosody Use in Speech Generation by Qwen2.5-Omni and Chroma

Determine whether Qwen2.5-Omni-7B and Chroma-4B leverage paralinguistic prosody when generating speech, rather than only when producing text outputs.

Background

The paper investigates how speaking style is encoded and used by audio-LLMs, but its output-level analysis is restricted to text because Whisper-large-v2 and Qwen2-Audio-7B-Instruct do not generate speech. Qwen2.5-Omni-7B and Chroma-4B do include speech-generation or voice-cloning components, yet the study does not evaluate whether those components preserve or exploit prosodic information during speech synthesis. Establishing this would extend the encoder-to-text analysis to the speech-generation pathway and clarify whether prosody is merely represented internally or actively used in generated audio.

References

We study only text outputs, since Whisper and Qwen2-Audio do not generate speech; whether Omni or Chroma additionally leverage prosody when generating speech is an open question.

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models  (2609.00727 - Koduru et al., 1 Sep 2026) in Section 5.5, “Limitations” (Section~\ref{sec:limitations})