---
title: Continual Neural Topic Model (CoNTM)
url: https://www.emergentmind.com/topics/continual-neural-topic-model-contm
type: topic
---

# Continual Neural Topic Model (CoNTM)

Searching arXiv for the target paper and closely related continual topic modeling work.
Continual Neural Topic Model (CoNTM) is a continual-learning topic model for time-sliced corpora in which the objective is to learn new topic models at subsequent time steps without forgetting previously learned topics. It is positioned between two earlier paradigms: Dynamic Topic Models, which learn topic evolution from the entire training corpus at once, and Online Topic Models, which are updated continuously from new data but do not have long-term memory. CoNTM addresses this gap by maintaining a global prior distribution that is continuously updated and from which each time-specific local topic model is derived [2508.15612].

## 1. Problem setting and conceptual placement

In the continual-learning formulation adopted by CoNTM, the standard objective of learning a new task without forgetting what was learned previously is translated into topic modeling as learning new topic models without forgetting previously learned topics. The model processes a sequence of time slices \(t=1,\dots,T\), with documents at each slice treated as the current task. The temporal coupling is not implemented by retraining on the entire historical corpus at every step; instead, the only connection across time is a shared global topic matrix that acts as long-term memory [2508.15612].

This placement is important because CoNTM is neither a conventional Dynamic Topic Model nor a purely online topic model. Dynamic Topic Models require access to the full corpus and model topic evolution in batch form, whereas online methods update incrementally but, as described in the source material, do not have long-term memory. A common misconception is that online updating alone is sufficient for continual topic modeling. CoNTM explicitly rejects that equivalence: its continual behavior depends on preserving an accumulated global prior so that older topics remain available when later slices are learned [2508.15612].

The model is designed for time-sliced data and is intended to capture both stable themes and emergent lexical shifts. This suggests a view of continual topic modeling in which temporal adaptation and memory preservation are treated as jointly constrained objectives rather than as separate post hoc corrections.

## 2. Architecture and topic parameterization

CoNTM builds on the Dirichlet VAE (DVAE) topic model. At each time step \(t\), it contains a local topic matrix \(\phi^{\loc}_t \in \mathbb{R}^{K\times V}\), with one row per topic and one column per vocabulary item, together with a global topic matrix \(\phi^{\glo} \in \mathbb{R}^{K\times V}\) that serves as long-term memory. The model also includes an encoder network \(q_\theta(z\mid w)\) that maps a bag-of-words document \(w\) to a Dirichlet approximate posterior over a topic mixture \(z \in \Delta^{K-1}\), and a decoder that mixes topics with \(z\) and applies a softmax over the vocabulary [2508.15612].

The key structural relation is that local topics are anchored to the global topics through an additive perturbation:
\[
\phi^{\loc}_t
=
g\bigl(\phi^{\glo},\Delta\phi^{\loc}_t\bigr)
=
\phi^{\glo}+\Delta\phi^{\loc}_t,
\qquad
\sum_{t=1}^T\Delta\phi^{\loc}_t=0.
\]

This formulation separates stable and time-specific components. The global matrix \(\phi^{\glo}\) functions as the persistent cross-slice representation, while \(\Delta\phi^{\loc}_t\) captures the deviation needed for slice \(t\). The source material states that the additive perturbation ensures local topics deviate only modestly from the global anchor, thereby promoting diversity between topics slice by slice. It also states that continual accumulation in \(\phi^{\glo}\) generates a “backbone” of stable themes, while \(\Delta\phi^{\loc}_t\) captures novel, emerging words [2508.15612].

Graphically, each document \(w_n\) at time \(t\) is generated from a latent topic proportion vector \(z\) with Dirichlet prior parameterized by \(\alpha\). The fact that the only cross-time dependency is mediated by \(\phi^{\glo}\) is central: memory is concentrated in topic space rather than in retained raw data or separate per-task parameter snapshots.

## 3. Generative process, inference, and training objective

For a document \(\mathbf w=(w_1,\dots,w_N)\) in slice \(t\), CoNTM uses the following generative process. First, draw the topic proportion vector
\[
z \sim \Dir(\alpha).
\]
Then, for each token \(n=1,\dots,N\),
\[
w_n \sim \Multi\bigl(1,\;\sigma\bigl(\phi^{\loc}_t\,z\bigr)\bigr),
\]
where \(\sigma\) is the row-wise softmax yielding a \(V\)-dimensional word distribution [2508.15612].

The true posterior \(p(z\mid w)\) is intractable, so CoNTM uses a variational approximation
\[
q_\theta(z\mid w)=\Dir\bigl(\alpha_\theta(w)\bigr),
\]
where \(\alpha_\theta(w)\) is produced by the encoder network. Training maximizes the usual evidence lower bound (ELBO) for each document:
\[
\mathcal L(\theta,\phi^{\glo},\Delta\phi_t^{\loc}\mid w)
=
-\,\KL\bigl(q_\theta(z)\,\|\,p(z\!\mid\!\alpha)\bigr)
+
\E_{q_\theta(z)}\bigl[\log p(w\mid z,\phi^{\glo},\Delta\phi_t^{\loc})\bigr].
\]

All parameters \(\theta\), \(\Delta\phi_t^{\loc}\), and indirectly \(\phi^{\glo}\), are updated by gradient ascent on the sum of ELBOs over the documents in slice \(t\) [2508.15612].

The formal consequence of this design is that temporal adaptation occurs through the learned perturbation \(\Delta\phi_t^{\loc}\), while long-term retention is mediated by the evolving global prior. This suggests a modular decomposition of temporal topic modeling into amortized inference over document-level mixtures and continual accumulation in a shared topic memory.

## 4. Global-prior updating and memory preservation

The continual-learning mechanism in CoNTM is implemented after fitting each slice. Once slice \(t\) has been learned, the local topics are formed as
\[
\hat\phi_t^{\loc}=\hat\phi^{\glo}+\Delta\hat\phi_t^{\loc},
\]
and the global topics are updated by a running average:
\[
\hat\phi^{\glo}\leftarrow(1-\rho_t)\,\hat\phi^{\glo}+\rho_t\,\hat\phi_t^{\loc},
\qquad
\rho_t\in(0,1).
\]
A typical schedule is
\[
\rho_t=1/(\tau_0+t)^\kappa,\qquad \kappa\in(0.5,1].
\]
In the reported experiments, the continual-learning schedule is
\[
\rho_t=1/(1+t)^{0.7},
\]
with sensitivity analyses showing robustness [2508.15612].

The source material gives an explicit interpretation of this schedule. Early slices have higher \(\rho\), so the global topics adapt rapidly; later slices influence the global topics more slowly, which prevents forgetting. Because each local model depends on \(\phi^{\glo}\), previously learned topics remain available as anchor points for new slices. At slice \(t\), only documents from that slice and the current \(\phi^{\glo}\) are needed. The global topics accumulate information from all prior slices through the weighted update, and this shared global prior ensures that old topics are never fully overwritten, eliminating catastrophic forgetting [2508.15612].

The model therefore treats memory preservation as a property of the update rule itself rather than as an auxiliary replay or distillation mechanism. The source material further states that balance in \(\rho_t\) prevents either rapid drift, which kills old topics, or overly rigid priors, which cannot accommodate new themes.

## 5. Experimental protocol and reported performance

The reported evaluation covers six datasets: New York Times (NYT, 1987–2007), UN General Debate transcripts (1970–2020), NIPS proceedings (1987–2019), NASA Tweets (2018–2022), arXiv titles+abstracts (2012–2024), and DBLP computer-science abstracts (2000–2020). The baselines are DETM (Dynamic Embedded Topic Model), DLDA (Blei & Lafferty’s Dynamic LDA), Dynamic BERTopic (clustering + c-TF-IDF), and DNLDA (Dynamic Noiseless LDA). The common hyperparameters are 50 topics, an 80/10/10 train/val/test split, and the Adam optimizer [2508.15612].

The evaluation metrics are topic coherence (TC), defined as NPMI against a temporal reference corpus; topic diversity (TD), defined as one minus the topic redundancy, namely average word overlap across topics; topic quality (TQ), defined as the normalized product \( \mathrm{TC}\times \mathrm{TD}\) across slices; temporal topic smoothness (TTS), which measures how abrupt or gradual topic shifts are; and predictive perplexity (PPL), defined as the log-likelihood of slice \(t+1\) under the model trained on slice \(t\), with lower values preferred [2508.15612].

The reported quantitative findings are specific. CoNTM attains the best average rank, \(1.6\), in TQ across the six corpora, outperforming DETM, DLDA, Dynamic BERTopic, and DNLDA. On the large corpora—NYT, DBLP, and arXiv—CoNTM occupies the upper-right corner of the coherence–diversity plot, indicating both high TC and high TD. Its predictive perplexity is the lowest on five of the six datasets, indicating strong generalization to new slices. Its TTS is approximately \(0.49\) on average, which is interpreted in the source material as meaning that topics evolve neither too abruptly nor remain static [2508.15612].

Qualitative analyses are also reported. In NYT, CoNTM detects the spike of “clinton” and “campaign” around 1992, whereas static-batch methods or purely online models either smear or lose that signal. In the UN debates, it smoothly shifts a “climate change” topic toward “Paris Agreement,” incorporating emerging terms such as “greenhouse,” “emission,” and “agreement” in 2015. The source material additionally states that topic coherence over time grows or remains stable in CoNTM, while Dynamic BERTopic often shows coherence degradation on larger, domain-specific corpora [2508.15612].

## 6. Relation to earlier lifelong neural topic modeling

CoNTM enters a line of research on continual or lifelong topic modeling but differs materially from earlier mechanisms. In “Lifelong Neural Topic Learning in Contextualized Autoregressive Topic Models of Language via Informative Transfers” [1909.13315], the base model is ctx-DocNADE, which combines DocNADE with an LSTM language-model component and augments it with three lifelong-learning tasks: Embedding Transfer (EmbTF), Selective Augmentation Learning (SAL), and Retention of Knowledge (RK). That framework is designed for sequential document collections and seeks to minimize catastrophic forgetting through informative transfers, selective replay, and soft parameter constraints [1909.13315].

A closely related framework is “Neural Topic Modeling with Continual Lifelong Learning” [2006.10909], which uses DocNADE as the backbone and combines topic regularization (TR), word-embedding guided transfer (EmbTF), and selective-data augmentation (SAL). In that formulation, past tasks contribute through TopicPool and WordPool, learned alignment matrices, and selective replay of historical documents whose current perplexity is sufficiently favorable. The paper reports improved perplexity, topic coherence, and information retrieval on sparse future tasks together with only minor degradation on earlier tasks [2006.10909].

Relative to those lifelong DocNADE-based approaches, CoNTM adopts a different memory mechanism. Rather than relying on TopicPool or WordPool accumulation, replayed documents, or projection-based alignment penalties, it maintains a continually updated global topic prior \(\phi^{\glo}\) and learns only slice-specific perturbations \(\Delta\phi_t^{\loc}\) on top of that prior [2508.15612]. This structural choice aligns with the temporal setting emphasized by CoNTM: at each slice, only the current documents and the current global topics are required. A plausible implication is that CoNTM targets a more explicitly temporal notion of topic continuity, whereas the earlier lifelong neural topic models were formulated around streams of document collections and multi-source transfer.

The broader significance of CoNTM, as supported by the reported experiments, is that continual topic modeling can be implemented through a global topic memory that is updated incrementally yet remains stable enough to prevent forgetting. In the source material’s summary, this yields a neural-VAE framework that captures topic changes online, learns more diverse topics, and better captures temporal changes than existing methods [2508.15612].

Source: https://www.emergentmind.com/topics/continual-neural-topic-model-contm