UnMuted: Wastewater SARS-CoV-2 Lineage Model
- UnMuted is a wastewater-based latent factor model that infers SARS-CoV-2 lineages by identifying temporally consistent clusters of mutations.
- The approach jointly estimates mutation memberships and lineage abundances using probabilistic models such as Binomial NMF and spline-based temporal smoothing.
- Empirical results show that UnMuted recovers epidemic patterns by matching mutation frequency trends without relying on fixed clinical lineage definitions.
UnMuted is most directly the name of a set of methods for defining SARS-CoV-2 lineages from wastewater by identifying temporally consistent mutation clusters rather than importing fixed lineage definitions from clinical phylogenetics (Becker, 5 Aug 2025). Its central premise is that wastewater sequencing predominantly observes mutation frequencies, counts, and coverage, not assignable full genomes, so lineage structure must be inferred from mixtures of mutations that rise and fall together over time. Related arXiv papers also use “UnMuted” as a broader interpretive lens for problems of recovering, verifying, or safely managing signals that are muted, missing, suppressed, or otherwise not directly accessible, but that broader usage is secondary to the formal wastewater method carrying the name (Wen et al., 13 May 2026, Liu et al., 1 Jul 2025, Yang et al., 2022).
1. Nomenclature and Scope
In the supplied literature, “UnMuted” has one explicit methodological meaning: the wastewater-lineage framework introduced in “UnMuted: Defining SARS-CoV-2 Lineages According to Temporally Consistent Mutation Clusters in Wastewater Samples” (Becker, 5 Aug 2025). That paper uses the term because “the mutations are allowed to speak for themselves,” emphasizing inference from wastewater mutation time series rather than from clinically curated lineage definitions.
Outside that usage, the label is not stable. “Muted: Multilingual Targeted Offensive Speech Identification and Visualization” explicitly states that it “does not mention ‘UnMuted’ anywhere,” and that from the paper alone “the correct system name is clearly Muted, not UnMuted” (Tillmann et al., 2023). Other works invoke “UnMuted” only as a query framing or interpretive contrast. “When Vision Speaks for Sound” treats “UnMuted” as a lens for asking whether a multimodal model can determine that a video is actually silent, sounded, synchronized, or acoustically trustworthy (Wen et al., 13 May 2026). This suggests that, across fields, “UnMuted” functions both as a proper method name and as a shorthand for recovering or validating latent signal content.
2. Wastewater Lineage Inference Problem
The UnMuted wastewater formulation begins from a specific surveillance constraint: clinical lineage assignment is based on near-complete genomes and phylogenetic placement, whereas wastewater typically yields degraded RNA, short reads, mutation counts, and per-site coverage rather than recoverable whole genomes (Becker, 5 Aug 2025). In that setting, one generally cannot determine which mutations came from the same genome, so the observable object is a mixed mutation-frequency signal.
The paper adopts the standard wastewater deconvolution principle that a mutation’s frequency is approximately the sum of the abundances of the lineages containing that mutation. In scalar form, the model assumes
where is the observed frequency of mutation at time , indicates whether mutation belongs to latent lineage , and is the abundance of lineage at time (Becker, 5 Aug 2025). In matrix form, the same idea is written as
0
The methodological departure of UnMuted is that 1 is not assumed known. Instead, the paper hypothesizes that “collections of mutations that appear together over time constitute lineages,” so lineage definitions and lineage abundances must be estimated jointly from the wastewater mutation time series (Becker, 5 Aug 2025). This replaces clinically fixed mutation barcodes with temporally inferred mutation clusters. A plausible implication is that the method is especially relevant when clinical definitions are incomplete, delayed, or inconsistent with wastewater-specific diversity.
3. Statistical Formulation and Model Family
UnMuted considers three main probabilistic models, together with a plain NMF baseline. All are variants of latent factorization, but they differ in whether they use read depth explicitly, whether lineage definitions are binary, and whether temporal smoothness is imposed (Becker, 5 Aug 2025).
The observation model in the binomial formulations is
2
where 3 is the count of reads supporting mutation 4 at time 5 and 6 is the read depth (Becker, 5 Aug 2025). The model constrains 7, and the abundance vectors are either forced to sum exactly to one or allowed to sum to at most one, depending on the variant.
| Model | Core constraint | Temporal structure |
|---|---|---|
| Binomial NMF | Exact compositional abundance via Dirichlet prior | None |
| TBNMF8 | Sum-to-one-or-less after deterministic normalization | B-spline abundance curves |
| TBNMF9 | Exact sum-to-one via Dirichlet over spline-derived concentrations | B-spline abundance curves |
For the non-temporal Binomial NMF, the paper specifies
0
with
1
(Becker, 5 Aug 2025). This yields binary mutation-lineage memberships and exact per-time compositional abundances, but no temporal smoothing.
The temporal models replace independent per-time abundances with spline-based trajectories. In TBNMF2, the abundance process is constructed from B-spline basis functions
3
with
4
then post-processed by a “fluffmaxing” function that leaves abundances unchanged if their sum is at most one and rescales them if the sum exceeds one (Becker, 5 Aug 2025). The paper’s prose makes clear that this is intended to allow missing or unmodeled lineages by enforcing a sum-to-one-or-less constraint rather than exact compositional closure.
In TBNMF5, the spline output serves instead as a Dirichlet concentration parameter:
6
so the abundances always sum to one, while the magnitude of the spline values controls compositional variability (Becker, 5 Aug 2025). The paper presents this as a smoother yet stochastic alternative to deterministic normalization.
4. Data, Estimation, and Empirical Behavior
The main application uses public wastewater data from NCBI BioProject PRJNA1088471, focusing especially on Highland Creek wastewater treatment plant samples because they were collected at a wastewater treatment plant rather than an airport site and were mostly equally spaced in time (Becker, 5 Aug 2025). The study reports 82 unique sampling dates, usually separated by 7 days.
Mutation preprocessing is substantive. The pipeline retains mutations with read depth at least 40 in a sample, then filters for temporal variation by requiring a mutation to be at least 10% frequency for 7 time points and less than 90% frequency for 8 time points, with 9. This yielded 88 mutations for 0, 72 for 1, and 57 for 2 (Becker, 5 Aug 2025). If multiple samples occurred on the same date, counts and coverages were summed before recalculating frequencies.
The Bayesian models were fit in NIMBLE in R, while rank selection for the probabilistic factorizations was guided by WAIC (Becker, 5 Aug 2025). The paper reports that no single rank diagnostic was definitive. For Binomial NMF, WAIC showed local minima at 10 and 13, but the paper ultimately discusses a 9-lineage model because rank 10 had one lineage set to 0. For TBNMF3, WAIC showed local minima at 6, 7, and 11, and the paper highlights a 7-lineage solution because higher-rank fits contained one or two always-zero lineages (Becker, 5 Aug 2025). For the temporal models, the number of B-spline basis functions was set to 10 after visual comparison of 8, 10, 12, and 14 basis functions.
Empirically, the independent GLM baseline ProVoC, using known lineages, recovered a sequence of waves at Highland Creek: BA.1, then BA.2.75, then BF.1, then BQ.1, then XBB.1.5, with XBB.1.9 still rising at the end of the period (Becker, 5 Aug 2025). UnMuted’s latent-factor models did not recover these as one-to-one named components, but they did recover temporally similar abundance patterns without clinical labels. Plain NMF already produced broad waves concordant with ProVoC. Binomial NMF often split known lineages into several inferred clusters, and one estimated lineage was not particularly similar to any barcode lineage and only gained abundance toward the end, which the paper interprets as either a superfluous factor or a possible cryptic lineage signal (Becker, 5 Aug 2025).
The temporal models are the strongest demonstration of the UnMuted premise. TBNMF4 and TBNMF5 recovered mutation clusters whose abundance trajectories matched known epidemic replacement patterns even when their mutation memberships were not exact barcode replicas. The paper repeatedly emphasizes that substantial overlap among BA.2.75-, BF.1-, BQ.1-, XBB.1.5-, and XBB.1.9-related mutations makes exact one-to-one recovery unrealistic, yet temporally consistent mutation clusters still behave like lineage signals (Becker, 5 Aug 2025).
5. Interpretation, Validation, and Limitations
UnMuted validates inferred lineages indirectly rather than through reconstructed full genomes. The main criteria are temporal plausibility, Jaccard similarity to UShER barcodes and PANGO constellations, cross-method consistency, and recurrence across multiple locations (Becker, 5 Aug 2025). A particularly important observation is that UShER barcodes and PANGO constellations themselves showed little similarity on the mutation subset used in the study, reinforcing the claim that “known” lineage definitions are not a uniquely stable gold standard.
The paper is explicit about ambiguity. Many mutations are shared across lineages, and among the selected Freyja lineages “there was no pair of lineages that did not share at least some mutations” (Becker, 5 Aug 2025). This makes exact identification underdetermined. Closely related or recombinant lineages can be split, merged, or partially re-expressed by latent factors that are nevertheless epidemiologically meaningful. The author therefore treats imperfect agreement with clinical definitions as expected rather than automatically erroneous.
Several limitations are concrete. Rank selection remains subjective even with WAIC and NMF diagnostics. The lineage-removal indicator variables did not reliably eliminate superfluous lineages, because the model still estimated temporal trends that could later be multiplied by zero, and the paper proposes RJMCMC as a possible alternative (Becker, 5 Aug 2025). Mutation preprocessing thresholds were acknowledged as somewhat arbitrary. Multi-location modeling, alternative temporal priors such as penalized splines, Gaussian processes, or autoregressive models, and stronger regularization of 6 to produce sparser lineage definitions are all identified as future extensions (Becker, 5 Aug 2025).
Taken together, these points place UnMuted less as a replacement for phylogenetic nomenclature than as a wastewater-native latent-structure model. It estimates lineage-like temporal components from wastewater alone, then compares them to external nomenclatures rather than assuming those nomenclatures are fully correct a priori.
6. Broader Research Uses of “UnMuted”
The surrounding literature suggests that “UnMuted” has become a recurring lens for signal-verification and signal-recovery problems beyond wastewater, although those works generally do not introduce a formal method with that name.
| Domain | Relation to “UnMuted” | Paper |
|---|---|---|
| Wastewater genomics | Official method name for temporal mutation-cluster lineage inference | (Becker, 5 Aug 2025) |
| Video MLLM audio grounding | Lens for verifying sound existence, synchronization, and consistency under Mute/Shift/Swap interventions | (Wen et al., 13 May 2026) |
| Silent-video speech generation | “Directly relevant” framing for generating speech from silent talking-face video | (Liu et al., 1 Jul 2025) |
| Video-conference privacy | Query about whether mute stops microphone access or only meeting transmission | (Yang et al., 2022) |
| Sensory accessibility | Framing for safe participation in live audio spaces under misophonia | (Ammari et al., 19 Jan 2026) |
| MARL communication | Converse of muting low-value messages: preserve only return-relevant communication | (Zuo et al., 3 Jul 2026) |
In multimodal video-language modeling, “When Vision Speaks for Sound” argues that many systems appear audio-grounded while actually relying on visual priors, and treats “UnMuted” as the capability to distinguish events that are merely visually associated with sound from clips that actually contain audible evidence (Wen et al., 13 May 2026). In silent-video synthesis, “MuteSwap” defines Silent Face-based Voice Conversion as generating speech from a silent source video plus target face images, which the paper describes as directly relevant to an “UnMuted” problem setting because it synthesizes plausible speech without any source audio (Liu et al., 1 Jul 2025). In privacy analysis, “Are You Really Muted?” reframes the issue as whether in-app mute truly stops microphone access; the paper shows that, for many native desktop video-conferencing apps, mute often prevents meeting participants from hearing audio without revoking microphone access, and in one case audio-derived telemetry enabled background-activity inference with a headline 81.9% macro accuracy (Yang et al., 2022). In HCI accessibility, “Remote Triggers” treats being “unmuted” as exposure to microphone-amplified chewing, breathing, keyboard noise, and related audiovisual triggers, arguing for channel-specific controls and real-time filtering (Ammari et al., 19 Jan 2026). In cooperative MARL, “MUTE” interprets the “UnMuted” query as the converse of indiscriminate silencing: only messages with counterfactual value for joint return should remain effectively unmuted (Zuo et al., 3 Jul 2026).
A cautious synthesis is that, outside wastewater genomics, “UnMuted” functions less as a stable term of art than as a recurring conceptual problem: whether an absent, muted, suppressed, or privacy-sensitive signal can be inferred, validated, or selectively re-enabled without collapsing the relevant semantics, safety properties, or task performance.