---
title: 'SWUM: Dynamic Software Word Usage Model'
url: https://www.emergentmind.com/topics/software-word-usage-model-swum
type: topic
---

# SWUM: Dynamic Software Word Usage Model

Searching arXiv for papers on "Software Word Usage Model" and related SWUM terminology.
Search results for "Software Word Usage Model SWUM".
Software Word Usage Model (SWUM), as it is situated in the supplied arXiv literature, belongs to a conceptual family of word-usage frameworks in which words are treated as dynamical objects whose usage changes over time due to intrinsic and contextual forces, and whose semantics can be made observable either through temporal dynamics or through generated natural-language realizations such as definitions and usage examples. In this framing, SWUM is not presented as a single formally specified model; rather, it appears as a comparative point for two adjacent research programs: a dynamical theory of phase-coupled lexical oscillations in diachronic corpora and a neural framework for generating definitions and usages from word embeddings [2304.09292; 1912.05898].

## 1. Conceptual placement of SWUM

The most explicit placement of SWUM in the supplied literature occurs in the discussion of the phase-coupled delay model of word usage. That paper states that it does not mention the Software Word Usage Model (SWUM) explicitly, but is clearly in the same conceptual family as other word-usage dynamics models. It further states that, like SWUM-style frameworks, it treats words as dynamical objects whose usage changes over time due to intrinsic and contextual forces. The novelty of that work is identified as a move from single-word dynamics to coupled-community dynamics, using a logistic-delay oscillator for each word and a Kuramoto phase interaction to account for semantic synchronization [2304.09292].

This comparative placement distinguishes two levels of analysis. At one level, word usage is modeled as a temporal process with endogenous oscillatory structure and exogenous sociocultural modulation. At another, usage is modeled as a semantic realization problem in which an embedding is asked to generate an example sentence that demonstrates how a word is used in context. This suggests that SWUM, in the present evidentiary frame, is best understood as a broad usage-centered perspective rather than as a uniquely delimited formalism.

A common misconception is to equate word usage with a purely static lexical property. The supplied literature rejects that simplification in two ways. The dynamical model argues that individual words do not evolve independently, but as members of semantically coherent communities. The neural usage-modeling paper argues that meaning is not exhausted by dictionary-style definition, because a word’s meaning is also reflected in contextual usage.

## 2. Dynamical word usage as delayed logistic oscillation

The dynamical formulation starts from a Volterra-type integro-differential logistic model,
\[
\dot u = Ru\left[1-\frac{1}{k}\int_{-\infty}^t G(t-\tau)\,u(\tau)\,d\tau\right],
\]
where \(u(t)\) is the usage of a word, \(R\) is its growth rate, \(k\) is the carrying capacity or slowly varying trend imposed by sociocultural context, and \(G\) is a memory kernel weighting past usage. The empirical motivation is that word frequencies in large diachronic corpora are not purely monotonic trends: after separating a slow trend component from the residual oscillatory component, the authors find robust, approximately 14–16 year cycles in the residuals. They interpret these cycles through a fashion-like mechanism in which interest in a topic rises, usage increases, sustained use produces saturation, and then usage declines until the topic becomes attractive again later [2304.09292].

The paper uses the strong kernel
\[
G(\tau)=\frac{4\tau}{\bar\tau^2}e^{-2\tau/\bar\tau},
\]
whose effect is equivalent to a distributed delay with mean delay \(\bar\tau/2\) years. With this kernel, the integro-differential equation is reduced to the three-dimensional delay-chain system
\[
\left\{
\begin{array}{ll}
\dot u=Ru\,(1-w/k),\\[4pt]
\dot v=\frac{2}{\bar\tau}(u-v),\\[4pt]
\dot w=\frac{2}{\bar\tau}(v-w).
\end{array}
\right.
\]
Here \(v\) and \(w\) are auxiliary variables introduced by the chain trick to represent the delayed inhibition acting on \(u\).

This system has two equilibria: the origin and the positive equilibrium \((u,v,w)^*=k(1,1,1)\), which corresponds to steady usage tracking the external trend. As the delay increases, the equilibrium loses stability through a Hopf bifurcation at
\[
\bar\tau = \frac{4}{R},
\]
creating a stable oscillatory regime. The paper also identifies a lower boundary,
\[
\bar\tau = \frac{8}{27R},
\]
below which the oscillations become too strongly damped to matter. In the authors’ interpretation, this places the model in a parameter region where oscillatory behavior is natural and tunable by the delay.

## 3. Phase reduction and semantic-community synchronization

To build a phase description, the positive equilibrium \((k,k,k)\) is translated to the origin and the dynamics are linearized around it. In the region \(\bar\tau>8/(27R)\), the linearized system has one real negative eigenvalue and a complex-conjugate pair. By changing to the eigenvector basis and then to cylindrical coordinates in the complex subspace, the dynamics are written in terms of an amplitude \(r\), a phase \(\phi\), and a stable transverse variable \(z\). The resulting phase-map form is
\[
\begin{split}
\begin{pmatrix} \dot{r}\\ \dot{\phi}\\ \dot{z} \end{pmatrix}
=
R_z(\phi)
\begin{pmatrix}
\operatorname{Re}(\Lambda_1) & -\operatorname{Im}(\Lambda_1) & 0\\
\operatorname{Im}(\Lambda_1) & \operatorname{Re}(\Lambda_1) & 0\\
0 & 0 & \Lambda_3
\end{pmatrix}
\begin{pmatrix}
r\cos\phi\\
\sin\phi\\
z
\end{pmatrix}
+
R_z(\phi)
\begin{pmatrix}
\operatorname{Re}(nl_1)\\
\operatorname{Im}(nl_1/r)\\
nl_3
\end{pmatrix},
\end{split}
\]
where \(R_z(\phi)\) is the rotation matrix around the axis \(z=(1,-1,1)\), \(\Lambda_1\) and \(\Lambda_3\) are the complex and real eigenvalues of the linearized system, and \(nl_1,nl_3\) are the transformed nonlinear terms. The characteristic equation is
\[
\Lambda^3+\frac{4}{\bar\tau}\Lambda^2+\frac{4}{\bar\tau^2}\Lambda+\frac{4R}{\bar\tau^2}=0.
\]
The point of this transformation is that it exposes the phase \(\phi\) as the natural variable through which words in a semantic community can be coupled [2304.09292].

For a community \(\mathpzc{N}\) of \(N\) words, the coupled-community model assigns each word its own copy of the phase-reduced logistic oscillator and adds Kuramoto-type all-to-all coupling among the phases:
\[
\frac{\lambda}{N}\sum_{j\in \mathpzc{N}}\sin(\phi_j-\phi_i),
\]
where \(\lambda\) is a global coupling strength. The paper also defines a more general coupling for the full network, mixing topic-like keywords and oscillatory semantic fields:
\[
\frac{\lambda_k}{\kappa_{k_i}}\sum_{j\in \mathpzc{N}_{\,k}}\sin(\phi_j-\phi_i) + \frac{\lambda_t}{\kappa_{t_i}}\sum_{j\in \mathpzc{N}_{\,t}}\sin(\phi_j-\phi_i),
\]
where \(\mathpzc{N}_{\,k}\) and \(\mathpzc{N}_{\,t}\) denote the keyword and topic communities containing word \(i\), and \(\kappa_i\) is the degree of node \(i\) within the relevant community. The weights \(\lambda_k\) and \(\lambda_t\) can differ, and the paper reports that semantic fields and keywords are reproduced by different but related coupling strengths, with \(\lambda_t\) roughly twice \(\lambda_k\).

Empirically, semantic fields show stronger phase synchrony than keywords, with an average order parameter around \(\bar\rho\sim 0.5\), whereas keywords are less coherent, around \(\bar\rho\sim 0.35\). Shuffling the words across communities gives a baseline around \(\bar\rho\sim 0.2\). Simulations of isolated logistic units driven only by trends already increase coherence somewhat, but adding weak Kuramoto coupling reproduces the observed coherence levels much better. The same paper emphasizes that coherent oscillations are produced across English, Spanish, French, German, and Italian using a single global coupling scale, and that the result is robust to network topology even when communities are sparsified from fully connected to \(N/2\), \(N/4\), or \(N/8\) random links.

## 4. Usage modeling as semantic realization

A second line of work models usage not as a temporal trajectory but as a conditional generation problem. In that formulation, the main goal is to improve the interpretability of word embeddings by making their semantic content observable as natural language. The paper does this in two related ways: definition modeling and usage modeling. Definition modeling asks a model to generate a dictionary-style definition from a target word embedding, while usage modeling asks it to generate an example sentence showing how the word is used in context. The paper treats these as complementary views of meaning: a definition explains what a word means, while a usage example shows how that meaning is realized in actual context [1912.05898].

Usage modeling is introduced to address a limitation of definition modeling. Definitions are abstract and lexical, but a word’s meaning is also reflected in contextual usage. The paper states: “We argue that the semantics captured by word embeddings can reconstruct its context in reverse. Thus, we introduce the usage modeling task to explore the possibility of using word embeddings to generate usages (example sentences).” It also emphasizes that usage modeling is harder than definition modeling because example sentences have many syntactic forms, tense variations, and lexical substitutions, whereas definitions have a more regular structure.

Formally, let the usage sentence be
\[
U=\left\{w_{1}, \dots, w_{N}\right\}
\]
with target word \(w^{*}\), and context
\[
C=\left\{c_{1}, \dots, c_{m}\right\}.
\]
The model estimates
\[
p\left(U | w^{*}, C\right)=\prod_{n=1}^{N} p\left(w_{n} | w_{i<n}, w^{*}, C\right).
\]
The target word should be included in the usage sentence, so the task is to generate a full sentence that uses the word appropriately given the target word and its contextual information. The paper explicitly treats usage modeling as a special case of language modeling, and states that its performance can be measured by perplexity on the test data.

## 5. Neural architectures for definition and usage generation

The single-task architecture is called Semantics-Generator. It uses a context encoder, context-aware sense attention, and a GRU decoder. The context encoder is a bidirectional GRU over the context \(C\), followed by max pooling to produce a context embedding \(\boldsymbol{v}_{c}\). The context-aware attention is scaled dot-product attention between the target word embedding \(\boldsymbol{v}^{*}\) and the context representations \(\boldsymbol{C}_{v}\), producing a context-aware target representation \(\boldsymbol{a}^{*}\). The decoder generates the output sequence autoregressively [1912.05898].

The attention and initialization equations are
\[
\boldsymbol{Q}=\boldsymbol{v}^{*} \boldsymbol{W}^{Q}, \boldsymbol{K}=\boldsymbol{C}_{v} \boldsymbol{W}^{K}, \boldsymbol{V}=\boldsymbol{C}_{v} \boldsymbol{W}^{V}
\]
\[
\operatorname{Attention}\left(\boldsymbol{Q}, \boldsymbol{K},\boldsymbol{V}\right)=\operatorname{softmax}\left(\frac{\boldsymbol{Q} \cdot \boldsymbol{K}^{\top}}{\sqrt{d}}\right) \boldsymbol{V}
\]
\[
\boldsymbol{a}^{*}=\boldsymbol{W}^{O} \cdot \operatorname{Attention}\left(\boldsymbol{Q}, \boldsymbol{K},\boldsymbol{V}\right)
\]
and
\[
\boldsymbol{s}_{0}=\boldsymbol{W}_{s}\left([\boldsymbol{v}^{*} ; \boldsymbol{v}_{c}]+\boldsymbol{b}_{s}\right).
\]
At each decoding step, the model uses a gated input,
\[
\boldsymbol{x}_{t}=\boldsymbol{g}_{t} \odot [\boldsymbol{a}^{*} ; \boldsymbol{y}_{t-1};\boldsymbol{c}^{*};\boldsymbol{e}^{*}],
\]
\[
\boldsymbol{g}_{t}=\sigma(\boldsymbol{W}_{g}  [\boldsymbol{a}^{*} ; \boldsymbol{y}_{t-1};\boldsymbol{c}^{*};\boldsymbol{e}^{*} ] ),
\]
\[
\boldsymbol{c}^{*}=\operatorname{CNN}(w^{*}), \qquad \boldsymbol{e}^{*}=\operatorname{ELMo}(C),
\]
followed by
\[
\boldsymbol{s}_{t}=\operatorname{GRU}\left(\boldsymbol{s}_{t-1},\boldsymbol{x}_{t} \right).
\]

The paper further proposes two multi-task models combining definition and usage generation. The Parallel-Shared Model shares the embedding layer and trains two separate decoders, one for definition generation and one for usage generation:
\[
\boldsymbol{s}_{t}^{d}=\operatorname{GRU}^{Def}\left(\boldsymbol{s}_{t-1}^{d},\boldsymbol{x}_{t}^{y}\right),
\]
\[
\boldsymbol{s}_{t}^{u}=\operatorname{GRU}^{Usg}\left(\boldsymbol{s}_{t-1}^{u},\boldsymbol{x}_{t}^{w}\right).
\]
The Hierarchical-Shared Model defines two versions, Hir-Shared-DU and Hir-Shared-UD. For Hir-Shared-DU, the usage decoder receives both the raw input and the lower-layer definition decoder output:
\[
\boldsymbol{s}_{t}^{\prime}=\operatorname{GRU}^{Def}\left(\boldsymbol{s}_{t-1}^{\prime},\boldsymbol{x}_{t}^{w}\right),
\]
\[
\boldsymbol{s}_{t}^{u}=\operatorname{GRU}^{Usg}\left(\boldsymbol{s}_{t-1}^{u},\boldsymbol{W}_{p} \left[\boldsymbol{x}_{t}^{w} ; \boldsymbol{s}_{t}^{\prime}\right]\right).
\]

For the definition model, the probability of a definition
\[
D=\left\{y_{1}, \dots, y_{T}\right\}
\]
is modeled as
\[
p\left(D | w^{*}, C\right)=\prod_{t=1}^{T} p\left(y_{t} | y_{i<t}, w^{*}, C\right),
\]
with
\[
p\left(y_{t} | y_{i<t},w^{*}, C \right)=\operatorname{softmax}\left(\boldsymbol{W}_{d} \boldsymbol{s}_{t}+\boldsymbol{b}_{d}\right).
\]
The paper states that the full definition model is trained by minimizing the negative log likelihood objective using Adam, and that in the multi-task setting “the sum of negative log likelihood objective of two tasks is optimized.”

## 6. Data resources, empirical results, and limitations

The usage-modeling paper uses two datasets. The original Oxford dataset from Gadetsky et al. (2018) contains 97,855 train, 12,232 valid, and 12,232 test examples, with 1,078,828 train tokens total across the train split. Its statistics include #Words of 33,128 / 8,867 / 8,850, definition length around 11, and context length around 17.7. The paper notes that this dataset has noisy entries and messy formatting, including target words containing Arabic numerals, non-English alphabets, and special symbols. To support usage modeling better, the authors collect Oxford-2019 using the Oxford Dictionaries API. It is built from the most common 65,000 tokens in WikiText-103; duplicates, function words, and stop words are removed; only pure-letter words are kept; and each specific definition includes three different context sentences and one usage sentence. The split is by specific meaning of words, so the same definition for a given word does not appear across train, valid, and test. Its statistics are #Words 32,066 / 7,359 / 7,322, #Entries 234,949 / 29,275 / 29,241, and #Tokens 2,595,561 / 323,700 / 322,915, with average context length about 21.4 [1912.05898].

For generation tasks, the paper uses BLEU and ROUGE-L, and for usage modeling specifically, perplexity. The usage-side results reported in Table 7 are 258.33 / 12.96 / 14.04 for Semantics-Generator, 237.64 / 13.49 / 14.32 for Parallel-Shared*, 227.71 / 13.51 / 14.47 for Hir-Shared-DU*, and 244.80 / 13.26 / 14.31 for Hir-Shared-UD*. The paper interprets these results as evidence that multi-task learning improves usage generation, that Hir-Shared-DU gives the best usage results among the listed models, and that definition and usage are beneficial to one another. It also states that this is the first study trying to generate example sentences by neural networks.

Qualitative examples show that the model can generate understandable, context-appropriate usages for different senses of the same word, including examples for **order** and **skirt**. At the same time, the paper explicitly notes that usage generation remains difficult, and that some outputs are still somewhat rough or grammatically odd, such as “She gaze at the room.” This supports a restrained interpretation: useful semantic structure is captured, but generation is not flawless.

The dynamical paper frames its model as a first step toward a universal dynamical theory of word usage. It is applied to large-scale Google Books data across English, Spanish, French, German, and Italian, focusing on nouns that persist over three centuries. It successfully explains both oscillatory cycles and within-community synchrony. However, it also notes substantial limitations: it analyzes only long-term dynamics and only nouns that remain in the corpus for the full period; it excludes words that enter or leave the corpus and finer-scale changes in community structure; it works at a coarse temporal scale; and it treats the community network as mostly static rather than modeling how communities themselves evolve [2304.09292].

Taken together, these results delimit the present evidentiary scope for SWUM. One neighboring literature models words as coupled oscillators poised near a Hopf bifurcation, where synchronization becomes possible through weak coupling. Another models usage as generated contextual realization that complements dictionary-style definition. A plausible implication is that SWUM, in this comparative setting, names a research orientation centered on usage as the primary observable of lexical structure, whether that observable is a long-range temporal trajectory or a context-conditioned sentence.

Source: https://www.emergentmind.com/topics/software-word-usage-model-swum