Frequency-Based Discourse Modulation
- Frequency-Based Discourse Modulation is a concept that examines measurable regularities in discourse through statistics, distributions, and spectral patterns of language signals.
- It integrates signal diversity, rate modulation, and harmonic regression techniques to reveal how coherence and information uptake are dynamically shaped across various contexts.
- Practical applications include modeling filled-pause rates in speech, adjusting discourse expectations in comprehension, and designing predictive models via spectral analysis in dialogue.
to=arxiv_search.search тәшкиqatigiit 凤凰大参考 code պարզjson {"query":"Frequency-Based Discourse Modulation discourse markers filled pauses oscillations harmonic information contours opinion shifts deliberative discourse", "max_results": 10} to=arxiv_search ашәа code պարզjson {"query":"Unpacking Ambiguity: The Interaction of Polysemous Discourse Markers and Non-DM Signals", "max_results": 5} Frequency-Based Discourse Modulation (FBDM) denotes a family of approaches in which discourse is shaped, interpreted, or modeled through measurable regularities in frequency. Across recent work, “frequency” is operationalized in several distinct but related ways: as distributions over discourse-marker senses and co-occurring signals, as event rates such as filled pauses per unit time, as base-rate priors for speaker-content expectations, as periodic structure in token-level surprisal contours, as recurrence frequencies of discussion points, as context-conditioned distributions of discourse relations, and as spectral structure extracted from discourse embeddings for predictive modeling (Wu et al., 22 Jul 2025, Porupski et al., 7 Jul 2026, Wu et al., 3 Feb 2025, Tsipidi et al., 4 Jun 2025, Tan et al., 2016, Cortez et al., 2023, Thakur et al., 26 Sep 2025). The unifying idea is that discourse behavior is not exhausted by categorical labels alone: distributions, rates, and oscillatory regularities act as constraints that modulate coherence, uptake, adaptation, and prediction.
1. Conceptual scope
The term has no single canonical definition across the literature. In work on discourse signaling, it refers to how the frequencies, distributions, and co-occurrence patterns of discourse relation signals—both discourse markers (DMs) and non-DM signals—shape interpretation in context (Wu et al., 22 Jul 2025). In parliamentary speech, it refers explicitly to statistical occurrence rates of filled pauses per unit time and not to acoustic spectral frequency (Porupski et al., 7 Jul 2026). In psycholinguistics, it refers to probabilistic adaptation driven by the base rate that a speaker will produce stereotype-incongruent content (Wu et al., 3 Feb 2025). In information-theoretic work, it refers to periodicity in information rate, modeled as oscillations in token-level surprisal across discourse units (Tsipidi et al., 4 Jun 2025). In computational deliberation, it denotes a frequency-spectrum fusion module that applies FFT-based modulation to discourse representations (Thakur et al., 26 Sep 2025).
| Domain | Frequency notion | Representative mechanism |
|---|---|---|
| Discourse signaling | Sense distributions and co-occurrence distributions | DM entropy and non-DM diversity |
| Speech production | Event rate per unit time | Negative Binomial GEE for FP rate |
| Comprehension | Prior probability / base rate | Speaker-general and speaker-specific updating |
| Information structure | Periodic oscillation in surprisal | Harmonic regression with time scaling |
| Deliberation modeling | Spectral frequency of embedding sequences | FFT/MLP/IFFT gating |
This multiplicity is not merely terminological. It suggests that FBDM is best understood as an umbrella over several distribution-sensitive research programs rather than a single theory with one metric. A common theme is that discourse interpretation depends on structured regularities beyond isolated words: signal diversity, rate modulation, prior probabilities, and periodic organization all function as modulators of discourse processing.
2. Distributional modulation in discourse signaling
A detailed formalization appears in work on polysemous discourse markers within the Enhanced Rhetorical Structure Theory (eRST) framework (Wu et al., 22 Jul 2025). eRST extends classic RST relation labels such as adversative-contrast, concession, causal-result, temporal-sequence, elaboration-attribute, attribution, question, and list with explicit signal annotations, and distinguishes DMs from non-DM signals across eight major classes and 45 subtypes: dm, graphical, lexical, morphological, numerical, reference, semantic, and syntactic. The framework supports concurrent signaling, multiple relations per node, and “tree breaking” edges.
Within this formulation, DM polysemy is graded rather than categorical. A DM’s polysemy is quantified by Shannon entropy over its eRST relation distribution:
where ranges over the eRST relations a DM can signal. For within-genre comparison, entropy is normalized as
Higher entropy indicates a more even sense distribution and therefore greater graded polysemy. The paper’s illustrative calculations show this logic directly: “since” with causal $0.6$ and temporal-circumstance $0.4$ yields bits; “but” with adversative-contrast $0.9$ and concession $0.1$ yields bits; “then” with sequence $0.8$ and result/inference 0 yields 1 bits. These examples are explicitly illustrative rather than reported corpus counts for specific DMs.
The corresponding claim about modulation is that higher DM entropy lowers the informativeness of the DM by itself and increases reliance on non-DM signals. Non-DM signals include graphical cues such as colon, dash, semicolon, quotation marks, headings, layout, and list numbering; lexical cues such as indicative words, alternate expressions, and restatements; morphological cues such as tense and mood; numerical cues such as counts and enumerations; reference chains; semantic relations such as antonymy, meronymy, lexical chain, synonymy/repetition, negation, and attribution source; and syntactic cues such as infinitival clauses, relative clauses, parallel constructions, participial clauses, reported speech, and inversion. Diversity is measured as the number of unique co-occurring non-DM signal types and combinations, normalized by DM frequency:
2
The empirical basis is the GUM corpus: 255 documents, 250,409 tokens, 16 genres, and 21,435 eRST discourse relations with signals. DM–non-DM co-occurrences total 1,372 instances, or 6.4% of annotated relations. By genre, the proportions are higher in essay (8.7%), biography (8.6%), and wikiHow (8.5%), and lowest in conversation (3.9%). Signal combination sizes are highly skewed toward small combinations: DM+1 signal accounts for 79.55%, DM+2 for 16.7%, DM+3 for 3.1%, DM+4 for 0.44%, and DM+5/6/8 each for 0.07%. The most frequent DM in co-occurrence is and at 36.6%; in academic writing, by is prominent in means relations such as “by using…”.
The statistical findings are specific. The Pearson correlation between DM entropy and the total number of co-occurring non-DM signals is weak but significant, 3, 4. More importantly, regression on normalized diversity shows that entropy predicts diversity whereas raw signal count does not. In Model 1, 5, the entropy coefficient is 6 (SE 7, 8), with 9 and adjusted 0. In Model 2, adding total signal count leaves entropy significant at 1 (SE 2, 3) while total signals is non-significant at 4 (SE 5, 6). In Model 3, genre moderates the entropy-to-diversity effect: the interaction is positive for vlog (7, SE 8, 9) and negative for letter ($0.6$0, SE $0.6$1, $0.6$2). A Chi-Squared Goodness of Fit test with FDR correction finds that all 16 genres deviate significantly from the global signal-type distribution for the same set of DMs ($0.6$3); spoken genres such as vlog and conversation favor reference signals, whereas written genres show more graphical and layout cues.
Concrete examples illustrate the mechanism. And often co-occurs with semantic lexical chains and meronymy; in conversation, personal reference chains accompany it to signal elaboration or listing. As in letters is strongly biased toward mode relations, approximately 65% versus 32.2% elsewhere, reducing ambiguity through genre priors. By in academic writing co-occurs with lexical signals such as “using,” clearly signaling means. This supports the paper’s central formulation: polysemous DMs co-occur with more diverse non-DM signals, but not necessarily with more signals overall.
3. Rate modulation in speech production
A second major formulation treats discourse modulation as variation in filled-pause frequency in spontaneous institutional speech (Porupski et al., 7 Jul 2026). The study operationalizes filled pauses as non-lexical vocalizations of the uh/um family detected directly from audio rather than from transcript tokens, including forms such as uh, um, mmm, and eh/emm-like vocalisations. The corpus is the ParlaSpeech Collection for Croatian, Czech, Polish, and Serbian parliamentary speech, with aligned official transcripts, speaker metadata, and utterance-level sentiment scores from XLM-R-ParlaSent. After filtering for regular MPs, utterances of at least 3 seconds and at least 10 words, and speakers with at least 10 utterances, the final dataset contains 1,001,787 utterances, 1,561 speakers, and 3,889 hours.
Detection is performed with wav2vec2-bert, trained on Slovenian spoken data and tested cross-lingually. Prior work by Ljubešić et al. (2025) reported event-level $0.6$4–$0.6$5 cross-lingually; in this study the event-level $0.6$6 across all four languages. Speech rate is derived as syllabic vowel count divided by utterance duration. Baseline FP rates differ by parliament: Croatian 1.82 FPs/min, Serbian 1.38, Czech 2.91, and Polish 3.47.
The rate model is a Negative Binomial GEE with a duration offset:
$0.6$7
with speakers as clusters and an independent working correlation structure. Time-varying covariates are decomposed using a Mundlak correction into within-speaker deviations and between-speaker means. This yields population-average inference while separating state-like from trait-like effects.
The global effects are substantial. Male speakers have lower FP rates than female speakers, with IRR $0.6$8 or $0.6$9, but this effect is language-specific: it is significant in Croatian and Serbian and non-significant in Czech and Polish. Age shows a negative association globally, IRR $0.4$0 per decade. Speech rate has a strong inverse relation with FP rate: global IRR $0.4$1 per 1 SD, within-speaker IRR $0.4$2, and between-speaker IRR $0.4$3. Sentiment is positively associated with FP rate: the global baseline IRR is $0.4$4 per point on the 0–5 scale, the between-speaker IRR is $0.4$5, and the within-speaker effect is approximately $0.4$6–$0.4$7 across all four parliaments. Political orientation shows no global association between speakers, but within-speaker shifts toward more rightward affiliation are associated with higher FP rates in Croatia (IRR $0.4$8, $0.4$9) and Poland (IRR 0, 1). Power status also modulates rates: globally, opposition speakers have lower FP rates than governing speakers (IRR 2), with significant effects in Czech and Croatian.
The paper is explicit that this is “frequency” as occurrence rate rather than spectral frequency. The modulation claim is therefore rate-based: FP frequency adapts to affective, ideological, and role-related discourse conditions, and these adjustments differ when comparing a speaker to themselves versus to other speakers. This gives FBDM a production-side interpretation in which hesitation markers are not random noise but a systematically modulated component of discourse.
4. Probabilistic and neural modulation in comprehension
In psycholinguistic work, FBDM is instantiated as probabilistic adaptation of discourse comprehension based on base-rate expectations about speakers (Wu et al., 3 Feb 2025). The study examines whether listeners update their mental representations of speakers when hearing stereotype-congruent or stereotype-incongruent utterances in Mandarin. The critical manipulation is the base rate of incongruent utterances: low base rate makes incongruency surprising, high base rate makes it expected or habituated.
Experiment 1 involved 30 native Mandarin speakers and a 3 within-participant design crossing Congruency and Base rate. Experiment 2 involved a distinct group of 30 participants and dissociated base-rate information from the target speaker by introducing an alternate speaker whose fillers established the base-rate manipulation. EEG was recorded with 128 active sintered Ag/AgCl electrodes, and time-frequency analysis used complex Morlet wavelets from 2 to 45 Hz. No significant ERP effects were observed in the N400 or P600 windows in either experiment.
The crucial findings are oscillatory. In Experiment 1, a high-beta cluster at 21–30 Hz and 220–330 ms over a central–parietal distribution showed a significant Congruency × BaseRate interaction, 4, SE 5, 6, 7. Incongruent utterances decreased high-beta power under low base rate, 8, but increased it under high base rate, 9. A theta cluster at 4–6 Hz and 320–580 ms over a left posterior-central distribution also showed a Congruency × BaseRate interaction, $0.9$0, SE $0.9$1, $0.9$2, $0.9$3: incongruency decreased theta power under low base rate, $0.9$4, and increased theta power under high base rate, $0.9$5. Theta additionally interacted with Openness, $0.9$6, SE $0.9$7, $0.9$8, $0.9$9; less open participants trended toward theta increases to incongruency, whereas more open participants trended toward theta decreases. In Experiment 2, only the high-beta effect persisted: the Congruency × BaseRate interaction remained significant, $0.1$0, SE $0.1$1, $0.1$2, $0.1$3, but no significant theta clusters were found for the target speaker.
The interpretation advanced in the paper is dual-mechanistic. High-beta oscillations index a speaker-general adjustment of discourse expectations, whereas theta oscillations index speaker-specific model updating. In predictive-coding terms, beta is associated with maintaining the current predictive set or decreasing when the set must change; theta is associated with memory operations and model updating. The paper formalizes the adaptation process with Bayes’ rule:
$0.1$4
Here the base rate changes the prior $0.1$5 over hypotheses about a speaker’s propensity to produce incongruent content. A high base rate reduces surprise and sustains an adapted state; a low base rate increases surprise and induces revision. This extends FBDM from production and signaling into real-time neural dynamics of discourse comprehension.
5. Harmonic and spectral modulation
A distinct line of work treats discourse as a structured signal whose information rate oscillates across discourse units (Tsipidi et al., 4 Jun 2025). For a document $0.1$6, token-level surprisal is defined as $0.1$7, and the information contour is the sequence
$0.1$8
Against a strong reading of Uniform Information Density, the paper proposes the Harmonic Surprisal hypothesis: deviations from mean information rate are often periodic rather than noise. Harmonic regression models the contour as
$0.1$9
or, in the paper’s Eq. 6 parameterization,
0
Amplitude is defined as 1. The paper’s time-scaling extension replaces 2 with 3, where 4 is the length of the discourse unit containing token 5, thereby aligning oscillations to EDU, sentence, paragraph, or document spans.
The empirical setting spans English, Spanish, German, Dutch, Basque, and Brazilian Portuguese. Using 10-fold cross-validation, EDU-scaled models significantly reduce MSE relative to baseline in all six languages: English 9.91 to 9.46, Spanish 14.63 to 13.83, German 12.43 to 11.31, Dutch 9.32 to 8.73, Basque 9.00 to 8.67, and Brazilian Portuguese 9.62 to 9.07. EDU-level low-order harmonics dominate the amplitude spectrum, often in the 6–7 range. In German, for example, the mean surprisal immediately before a paragraph boundary is approximately 1.47 and after it 7.75; before a sentence boundary, 1.07 and after, 7.06; before an EDU boundary, 1.30 and after, 6.39. The reported pattern across all six languages is that surprisal dips before boundaries and rises just after them.
A different spectral formulation appears in deliberation modeling, where FBDM is a frequency-spectrum fusion module for predicting post-exposure opinions (Thakur et al., 26 Sep 2025). Here pre-survey answer sequences and presentation content are embedded with Sentence-BERT, transformed into the frequency domain, compressed with a small MLP over magnitude spectra, and reconstructed by inverse FFT into a modulation signal 8 that gates encoder hidden states:
9
The full pipeline uses a 4-layer, 4-head Transformer, AdamW with learning rate $0.8$0, weight decay $0.8$1, cosine annealing, and gradient clipping at 1.0. On a self-sourced dataset with 100+ university students and three topics—skincare products, ketchup, and DNA storage—the base Transformer achieves accuracy 0.757 and macro-F1 0.713; adding FBDM leaves accuracy at 0.757 but raises macro-F1 to 0.735. A Quantum-Deliberation Framework that prepends a simulated 2-qubit quantum token increases performance further to accuracy 0.878 and macro-F1 0.866. In this setting, FBDM is not a descriptive theory of natural discourse organization but a learned modulation mechanism that amplifies shared spectral patterns between respondent priors and presentation stimuli.
Taken together, these studies suggest two non-equivalent but compatible spectral conceptions of FBDM: one descriptive, in which discourse structure induces measurable oscillations in surprisal contours, and one engineering-oriented, in which spectral concordances are used to gate representations for downstream prediction.
6. Recurrence, dialogue distributions, and open problems
Frequency-based modulation also appears in studies of rhetorical uptake and dialogue-structure distributions. In analysis of Federal Open Market Committee transcripts from 1977 to 2008, discussion points are proxied by content words and tracked via matched in-context versus out-of-context word pairs within the same speech (Tan et al., 2016). Past frequency is defined as $0.8$2 and future recurrence is measured either within the same meeting or across the next five meetings. The effect measure is the recurrence advantage $0.8$3, where $0.8$4 is the average rate at which an in-context word recurs more than its matched control. The results are directionally specific: hedges such as “maybe,” “I think,” and “I don’t know” show a clear positive inter-meeting effect; superlatives show a negative effect; contrastive conjunctions such as “but” have little effect; and second-person pronouns “you” have a strong positive immediate intra-meeting effect that decays over time. The positive hedging effect is more pronounced for female speakers than for male speakers. This gives FBDM a persistence-based interpretation: rhetorical de-emphasis can increase the long-run recurrence frequency of discussion points.
In spontaneous conversation, discourse relations themselves exhibit context-conditioned distributions (Cortez et al., 2023). Using 19 Switchboard dialogues, 464 EDU/CDU pairs, and 114 novice annotators, the study adapts SDRT in the tradition of Asher and Lascarides and the STAC guidelines. Agreement is modest but above chance: soft-match adjusted $0.8$5, boot-match adjusted $0.8$6, and boot-F1 adjusted $0.8$7. Single-turn cases elicit more labels per pair and higher classifier cross-entropy than across-speaker cases, indicating greater uncertainty. Statistical modeling confirms that context matters: adding within/across-speaker and within/across-turn indicators improves fit by $0.8$8, $0.8$9, and adding annotator group/topic improves fit further by 00, 01. Across speakers, Acknowledgement, Clarification Question, Comment, and Question-Answer Pair are more likely; within a single speaker, Continuation, Elaboration, Explanation, and Narration are more likely. A ridge classifier over BERT next-sentence-prediction-style embeddings achieves strict annotation-level macro precision 0.21, recall 0.19, and F1 0.19, with top-guess recall of any annotator-provided label at 0.76 overall. This extends FBDM to dialogue management: discourse-relation priors shift with turn structure and speaker configuration.
The literature is also explicit about limitations. Entropy alone has modest explanatory power in DM polysemy modeling, with adjusted 02; local semantics and relation inventory still matter (Wu et al., 22 Jul 2025). Parliamentary FP results are domain-restricted, sentiment is automatically estimated with 03, and orientation effects depend on relatively rare party-switching episodes (Porupski et al., 7 Jul 2026). Neural adaptation results rely on stereotype-based Mandarin stimuli and text-to-speech voices, with null N400/P600 effects raising questions about task sensitivity (Wu et al., 3 Feb 2025). Harmonic surprisal depends on LM-derived surprisal and on corpora drawn from limited genres (Tsipidi et al., 4 Jun 2025). The FOMC recurrence study is observational and uses controlled matching rather than parametric regression (Tan et al., 2016). The deliberation modeling study uses a self-sourced dataset with 100+ students and reports no statistical significance testing for model comparisons (Thakur et al., 26 Sep 2025).
These constraints do not collapse the concept; they delimit it. The cumulative record suggests that FBDM is best treated as a research program centered on distribution-sensitive discourse analysis. Its most stable findings are that diversity of accompanying cues often matters more than raw count, that event rates can be socially and institutionally modulated, that comprehension adapts to base-rate priors at multiple representational levels, that information contours exhibit discourse-aligned periodicity, and that recurrence and dialogue-structure frequencies provide actionable priors for modeling. A plausible implication is that future work will continue to integrate these strands by combining genre-conditioned priors, signal diversity, temporal rate modeling, and spectral representations within a single account of discourse modulation.