---
title: 'FD-SLMs: Simultaneous Speech Interaction'
url: https://www.emergentmind.com/topics/full-duplex-speech-language-models-fd-slms
type: topic
---

# FD-SLMs: Simultaneous Speech Interaction

Full-Duplex Speech Language Models (FD-SLMs) are machine learning systems designed to enable simultaneous, low-latency spoken dialogue—allowing a single model to both listen to and emit speech concurrently, thus mirroring key aspects of natural human conversational dynamics such as overlapping speech, backchannels, context-dependent barge-in, and robust echo handling. Through recent architectural and algorithmic advances, FD-SLMs have established themselves as a foundational paradigm for human-computer interaction, outperforming modular, half-duplex predecessors in responsiveness, naturalness, and fluidity [2505.17060][2411.18138][2509.14515].

## 1. Defining Full-Duplex Speech Language Models

FD-SLMs formalize dialogue as a continuous, synchronous mapping between a stream of incoming audio (user speech and environment) and an outgoing stream of generated audio (machine response). The system must concurrently estimate both $P(s_t | x_{1:t}, s_{1:t-1})$ and $P(x_t|x_{1:t-1}, s_{1:t-1})$ at each timestep $t$, where $s_t$ is the system output and $x_t$ is the environmental input [2505.17060]. Unlike turn-based (half-duplex) systems, where agents defer speaking until a segmented utterance has been processed via ASR, FD-SLMs operate in a joint generative loop with tightly-coupled perception and verbalization pipelines.

Key dynamics of FD-SLMs include:
- **Overlap Handling:** Emitting speech concurrently with input, allowing both interruption (barge-in) and conversational backchannels (“mm-hm”, “right”) [2508.07375][2509.14515].
- **Echo Cancellation:** Avoiding feedback artifacts by dynamically modeling and gating self-generated speech when emitted to the environment [2411.18138].
- **Fine-grained Turn-Taking:** Learning when to yield or resume the conversational floor based on semantic and prosodic context.

## 2. Architectures and Synchronization Strategies

### 2.1 Engineered Synchronization (Modular Architectures)

Modular approaches rely on explicit duplex control modules (finite-state machines or external controllers) which arbitrate “speak”, “listen”, or “idle” states and explicitly gate the LLM’s generative process. Example realizations: FlexDuo (plug-in controller decoupled from LLM) uses semantic integrity buffering and sliding-window state machines for filtering and interruption handling [2502.13472]; Freeze-Omni and VITA-1.5 embed FSM or Voice Activity Detection (VAD)-driven arbitration as external mediators [2509.14515].

#### Table 1. Modular Synchronization Approaches

| Model         | Arbitration       | Data Flow                         |
|---------------|------------------|-----------------------------------|
| FlexDuo       | 3-state FSM      | ASR/NLU→FSM→LLM→TTS               |
| Freeze-Omni   | Internal tokens  | LLM (with Speak/Listen tokens)    |
| VITA-1.5      | External FSM     | FSM over two LLMs & VAD           |

*FlexDuo and VITA-1.5 exhibit predictable but sometimes delayed switching; their explicit structure is easily extensible but incurs latency and can propagate errors from imperfect perception modules [2502.13472][2509.14515].*

### 2.2 Learned Synchronization (End-to-End FD-SLMs)

End-to-end FD-SLMs internalize synchrony, jointly modeling user and agent streams via a single autoregressive backbone. Canonical instantiations:
- **Codecs-in-Token-Space:** Moshi, syncLLM, OmniFlatten tokenize audio via neural codecs (e.g., HuBERT, VQ-VAE) and interleave quantized speech tokens in an LLM [2509.14515][2410.17799], but this introduces significant modality gaps and re-training burdens.
- **Codec-Free Embeddings:** SALMONN-omni discards audio tokenization, instead using continuous log-Mel/embedding streams with cross-modal attention and a learned “thinking” mechanism for state transition [2505.17060][2411.18138].
- **Next-Token-Pair Prediction:** FLEXI and related work propose architectures where each step produces both the next output token and a control signal (e.g., “continue”, “yield”) via a single Transformer head, yielding lower latency and more precise arbitrations [2509.22243].

*End-to-end FD-SLMs discover conversational behaviors—overlap, barge-in, backchannel—as emergent properties, but require careful handling of temporal alignment and often massive multi-modal datasets for effective training [2409.15594][2509.14515].*

## 3. Training, State Transition, and Dynamic Control

### 3.1 Dynamic State Selection

Modern FD-SLMs such as SALMONN-omni deploy explicit state-transitions as special “thinking” tokens (⟨think⟩, ⟨shift⟩, ⟨start_speak⟩, ⟨end_speak⟩) within their token streams. Probability of transitioning to speak, listen, or think is estimated via a learned Bernoulli (sigmoid over the LLM’s hidden state and acoustic embeddings). Supervision is applied via cross-entropy loss over both ordinary and state tokens [2505.17060][2411.18138].

### 3.2 Control Tokenization and FSM Integration

Both engineered and hybrid models introduce control tokens or FSMs (e.g. [S.SPEAK], [C.LISTEN]) into the LLM vocabulary, sometimes informed through prompt engineering, instruction-tuning, or supervised annotation of dialogue states [2405.19487][2502.13472].

### 3.3 Reinforcement Learning for Turn Management

FD-SLMs applying reinforcement learning techniques (e.g., Direct Preference Optimization) further refine the timing of barge-in and backchannel behavior by maximizing reward for desirable real-time interruption handling [2505.17060].

## 4. Benchmarks and Evaluation

Recent work has established rigorous evaluation frameworks covering temporal, behavioral, semantic, and acoustic performance:

- **Temporal Dynamics:** Metrics include response latency (aiming for <200 ms), overlap ratio, and first-token offset (FTO) [2509.14515][2509.22243].
- **Behavioral Arbitration:** Assesses interruption response delay (IRD), barge-in success rate, and precise word error rate (WER) during arbitration [2507.23159][2509.22243].
- **Semantic Coherence:** Perplexity (PPL) and GPT-generated content scores capture the model’s ability to maintain meaningful dialogue [2509.14515][2410.17799].
- **Acoustic Quality:** Human (MOS) and automatic (UTMOSv2) ratings measure perceived naturalness and intelligibility [2507.23159].
- **Benchmark Suites:** Full-Duplex-Bench v1.5/v2, MTR-DuplexBench, FLEXI—test overlap handling, multi-turn coherence, barge-in, safety, and turn-taking across both open-source and commercial LLMs [2511.10262][2507.23159][2510.07838][2509.22243].

#### Table 2. Representative Full-Duplex Model Performance

| Model         | FTO (s) | Barge-in F1 | MOS | Unique Features                          |
|---------------|---------|-------------|-----|------------------------------------------|
| SALMONN-omni  | 0.38    | 0.88–0.93   |3.85| Standalone, codec-free, thinking tokens  |
| Moshi         | 2.22    | 0.80        |3.90| Codec injection, high data requirement   |
| FlexDuo       |  —      |  —          | —   | Modular FSM with semantic buffering      |
| Freeze-Omni   |  —      | 0.68        | —   | VAD-driven, two LLM processes           |

SALMONN-omni and SYNC-LLM consistently deliver lower latency (FTO < 0.4 s), higher barge-in/backchannel F1, and competitive MOS using substantially less training data compared to prior systems [2505.17060][2409.15594][2411.18138].

## 5. Empirical Findings, Limitations, and Comparative Analysis

Evaluations highlight common patterns:
- **Error Cascade in Modular Systems:** VAD and semantic integrity mismatches propagate interrupt errors and context pollution [2502.13472][2510.07838].
- **Modality Gaps in Codec-Injection:** Tokenizing audio as discrete codes requires large-scale speech-text alignment and can degrade intrinsic language ability [2505.17060][2509.14515].
- **Responsiveness vs. Robustness Tradeoffs:** Repair-first agents yield quickly in overlap settings but risk spurious interruption by noise or backchannels; continuity-first agents prioritize flow, risking delayed handover [2507.23159].

Benchmarks such as Full-Duplex-Bench and FLEXI reveal persistent weaknesses in multi-turn consistency, context drift, and handling of emergency or ambiguous overlap scenarios; even state-of-the-art systems lag behind human-level reactivity and semantic precision, especially under multi-round, noisy, or adversarial conditions [2511.10262][2509.22243][2510.07838].

## 6. Open Challenges and Future Directions

Open research fronts include:
- **Synchronous Data Scarcity:** There remains a lack of large, real human-human full-duplex speech corpora with annotated overlaps, interruptions, and nuanced timing [2509.14515]. Synthetic pipelines (TTS-driven, adversarially generated) are used to supplement these gaps but may not fully capture natural entrainment and topic flow.
- **Multi-party and Multimodal Extension:** Existing FD-SLMs mostly target dyadic (two-speaker) English conversation; scaling to multi-speaker (diarization), multi-modal (gestural, visual cues), and cross-lingual settings is largely unaddressed [2505.17060][2508.07375].
- **Hierarchical and Adaptive Modeling:** Prosody-aware objectives, explicit emotion and floor-control modeling, hierarchical reinforcement learning for complex goal management, and adaptive latency tuning present promising avenues for enhancing alignment with human discourse [2505.17060][2508.07375][2411.18138].
- **Benchmark and Protocol Standardization:** Continued development of open, streaming benchmarks, standardized task sets, and low-latency evaluation protocols (e.g., FDB v2, MTR-DuplexBench, FLEXI) are essential for reproducibility and fair comparison [2510.07838][2511.10262][2509.22243].

## 7. Representative and Milestone Models

A non-exhaustive list of significant FD-SLMs and their core contributions:
- **SALMONN-omni:** First codec-free, standalone model with dynamic thinking state, achieves 30%+ relative improvement on dialogue and barge-in metrics over prior art [2505.17060][2411.18138].
- **FlexDuo:** Modular plug-in controller achieving ~25% reduction in false interruptions; demonstrates that explicit buffering and semantic integrity control can retrofit half-duplex models [2502.13472].
- **OmniFlatten/FLM-Audio:** Native full-duplex via "flattened" token streams or natural monologue dual-training; demonstrates lower preprocessing cost and strong language fidelity [2410.17799][2509.02521].
- **LSLM:** Listening-While-Speaking fusion strategies (middle fusion) preserving TTS fidelity under simultaneous streaming inputs [2408.02622].
- **Synchronous LLM:** Explicit time embedding + scheduler produces human-like turn-taking, overlap, and backchannel patterns with minimal increases in generation latency [2409.15594].

In sum, FD-SLMs represent a convergence of advanced speech and language modeling toward truly synchronous, contextually-aware, fluid spoken interaction, with current research emphasizing convergence of modular and end-to-end paradigms, robust low-latency control, and holistic, multi-dimensional evaluation [2505.17060][2509.14515][2411.18138].

Source: https://www.emergentmind.com/topics/full-duplex-speech-language-models-fd-slms