---
title: 'K-Forcing: Concepts in Graph Theory and Beyond'
url: https://www.emergentmind.com/topics/k-forcing
type: topic
---

# K-Forcing: Concepts in Graph Theory and Beyond

K-Forcing is a context-dependent term rather than a single canonical construction. In graph theory, \(k\)-forcing is a propagation process on vertex colorings that generalizes zero forcing and yields the graph parameter \(F_k(G)\) or \(Z_k(G)\) when the literature uses that notation [1405.7573]. In model theory and set theory, the forcing notion \(\mathbb P_{\mathcal K,\kappa}\) builds a \(\kappa\)-sized structure from a Fraïssé class \(\mathcal K\) and recovers the classical Fraïssé limit when \(\kappa=\omega\) [1902.06206]. In extremal matrix theory, \(K\)-forcing and strongly \(K\)-forcing are pattern-enforcement properties for \((0,1)\)-matrices relative to a fixed template \(K\) [2510.27076]. In language modeling, K-Forcing denotes a push-forward language modeling paradigm for joint next-\(k\)-token decoding [2606.10820]. This suggests a shared metaphor—local admissibility conditions compel larger-scale structure—even though the underlying objects and techniques are different.

## 1. Graph-theoretic \(k\)-forcing: definition and core invariants

For a finite simple undirected graph \(G=(V,E)\), one colors a subset \(S\subseteq V\) black and leaves the remaining vertices white. Fix an integer \(k\ge 1\). The color-change rule is: whenever a black vertex \(v\) has at most \(k\) white neighbors, \(v\) \(k\)-forces each of those white neighbors to become black. If repeated application of the rule eventually colors all vertices black, then \(S\) is a \(k\)-forcing set, and the \(k\)-forcing number \(F_k(G)\) is the minimum cardinality of such a set. When \(k=1\), the process is exactly classical zero forcing, so \(F_1(G)=Z(G)\) [1405.7573].

The graph-theoretic literature records several basic structural facts. For the complete graph \(K_n\), \(F_k(K_n)=\max\{n-k,1\}\). There is a lower bound \(F_k(G)\ge \delta-k+1\), where \(\delta\) is the minimum degree, and the parameter is monotone in \(k\) in the sense that \(F_k(G)\ge F_{k+1}(G)\) [1401.6206]. These statements place \(k\)-forcing between minimum-degree constraints and a thresholded propagation dynamics.

Upper bounds were first developed in a degree-counting framework. If \(G\) has order \(n\ge 2\), maximum degree \(\Delta\ge k\), and minimum degree \(\delta\ge 1\), then
\[
F_k(G)\;\le\;
\frac{(\Delta-k+1)\,n}
{\Delta-k+1+\min\{\delta,k\}}.
\]
If \(\delta\ge k\), this simplifies to
\[
F_k(G)\;\le\;\frac{(\Delta-k+1)\,n}{\Delta+1},
\]
and for \(k=1\) one recovers
\[
Z(G)\;\le\;\frac{\Delta n}{\Delta+1}.
\]
For \(k\)-connected graphs with \(n>k\) and \(\Delta\ge 2\), a sharper bound is
\[
F_k(G)\;\le\;\frac{(\Delta-2)\,n+2}{\Delta+k-2}.
\]
The same paper also proves \(F_k(G)\le n-\gamma_{k,c}(G)\), where \(\gamma_{k,c}(G)\) is the connected \(k\)-domination number, and derives the zero-forcing corollary \(Z(G)+\gamma_c(G)\le n\) for connected graphs [1401.6206].

These results already tie \(k\)-forcing to domination, connectivity, and extremal degree structure. They also resolve a question posed by Meyer on regular bipartite circulant graphs, since the bound \(Z(G)\le \Delta n/(\Delta+1)\) applies in greater generality than the original question [1401.6206].

## 2. Dynamic and greedy formulations, refined bounds, and power-domination links

A later line of work replaces purely static edge-counting arguments by a dynamic greedy construction. For a connected graph \(G\), if \(\Delta\le k\), any single vertex is a \(k\)-forcing set. Otherwise one chooses a vertex \(v\) of minimum degree \(\delta\), colors \(v\) together with \(t=\max\{0,\deg(v)-k\}\) of its neighbors, runs the forcing process until it stalls, and whenever it stalls at a black vertex \(u\) with more than \(k\) white neighbors, colors exactly enough additional white neighbors of \(u\) so that \(u\) has at most \(k\) white neighbors and can force again. The resulting greedy set realizes the main upper bounds in the paper [1405.7573].

The corresponding theorem distinguishes three regimes. If \(\Delta\le k\), then \(F_k(G)=1\). If \(\Delta=k+1\) and \(\delta<\Delta\), then \(F_k(G)=1\); if \(\Delta=k+1\) and \(\delta=\Delta\), then \(F_k(G)=2\). When \(\Delta>k+2\), a useful simplified corollary is
\[
F_k(G)\le \frac{(\Delta-k-1)n+2k}{\Delta-1},
\]
with equality only if \(G\) is \((k+2)\)-regular. In the zero-forcing case this yields
\[
Z(G)\le \frac{(\Delta-2)n+2}{\Delta-1},
\]
and also
\[
Z(G)\le \frac{(\Delta-2)n-(\Delta-\delta)+2}{\Delta-1}.
\]
The note states that these corollaries improve two theorems from Amos, Caro, Dávila, and Pepper, and that the equality analysis sheds light on the regularity requirement in the equality case [1405.7573].

The same note explicitly observes that Meyer's question on bounding the zero-forcing number of bipartite circulant graphs in terms of \(\Delta\) and \(n\) is answered affirmatively, because
\[
Z(G)\le \frac{(\Delta-2)n+2}{\Delta-1}
\]
holds for any graph, not only for bipartite circulants [1405.7573].

A separate but closely related development studies the relationship between \(k\)-forcing and \(k\)-power domination. In \(k\)-power domination one starts from a power-dominating set \(S\), initially observes \(N[S]\), and thereafter applies the same propagation rule as in \(k\)-forcing. The paper establishes
\[
\gamma_{P,k}(G)\le Z_k(G)\le \gamma_{P,k}(G)\,(\Delta(G)+1),
\]
and strengthens the upper bound to
\[
Z_k(G)\le \gamma_{P,k}(G)\,(\Delta(G)+1-k)
\]
when \(\Delta(G)\ge k+2\), equivalently
\[
\gamma_{P,k}(G)\ge \frac{Z_k(G)}{\Delta(G)+1-k}.
\]
It also introduces contraction-based inequalities for \(G/X\) and the auxiliary graph \(\widehat X\), including a partition theorem of the form
\[
Z_k(G)\le \sum_{i=1}^r Z_k(\widehat P_i),
\]
provided each \(\widehat P_i\) has a minimum \(k\)-forcing set lying in \(P_i\). The stated motivation is parallel or divide-and-conquer computation of forcing sets on decomposable graphs [1701.08386].

## 3. Oriented and other graph-theoretic extensions

The oriented version replaces undirected adjacency by out-neighborhoods. For an orientation \(D\) of a simple graph \(G\), a colored vertex \(u\) with at most \(k\) uncolored out-neighbors forces each of those out-neighbors. The minimum size of a \(k\)-forcing set is \(F_k(D)\). Varying over all orientations of \(G\) gives the extremal invariants
\[
\MOF_k(G)=\max\{F_k(D):D\text{ is an orientation of }G\},
\qquad
\mof_k(G)=\min\{F_k(D):D\text{ is an orientation of }G\}.
\]
If \(T_k(G)\) denotes the minimum number of trees in a \((k+1)\)-tree cover of \(G\), then \(\mof_k(G)=T_k(G)\). For \(k=1\), this specializes to \(\mof(G)=\rho(G)\), the path-covering number. On the opposite side, \(\MOF_k(G)\ge \alpha(G)\), and equality holds if \(k\ge \Delta(G)\) or if \(G\) is a tree [1709.02988].

The oriented theory also admits degree-based estimates. If \(D\) has minimum out-degree \(\delta^+(D)\), then
\[
F_k(D)\ge \max\{\delta^+(D)-k+1,1\}.
\]
If \(D\) is reachable, has order \(n\), and maximum out-degree \(\Delta^+\), then
\[
F_k(D)\le \frac{(\Delta^+-k)n+k}{\Delta^+}.
\]
Examples include \(\mof(P_n)=1\), \(\MOF(P_n)=\lceil n/2\rceil\), \(\mof(C_n)=2\), and \(\MOF(C_n)=\lceil n/2\rceil\) when \(k=1\). For stars \(K_{1,n-1}\), one has \(\mof_k(K_{1,n-1})=\MOF_k(K_{1,n-1})=n-k-1\) for general \(k<n-1\) [1709.02988].

For complete graphs, where orientations are tournaments, the extremal quantity \(\MOF_k(K_n)\) becomes a tournament parameter. For \(k=1\), one paper proves two lower bounds:
\[
\MOF(K_n)\ge \frac{3}{4}n-O(1)\quad (n\ge 10),
\]
and, for all \(n\ge 2\),
\[
\MOF(K_n)\ge n-\frac{2n}{\log_2 n}.
\]
For general \(k\), the transitive tournament satisfies
\[
F_k(D)=\left\lceil \frac{n}{k+1}\right\rceil,
\]
so
\[
\MOF_k(K_n)\ge \left\lceil \frac{n}{k+1}\right\rceil.
\]
The same paper gives multipartite lower bounds such as
\[
\MOF_k(G)\ge n_1+\sum_{i=2}^q \max\{n_i-k,0\}
\]
for complete \(q\)-partite graphs with part sizes \(n_1\ge \cdots \ge n_q\) [1709.07509].

A distinct use of forcing terminology appears in the study of local majority on connected graphs. There, for an infinite family \(\mathcal G\), an edge weighting \(w:E(G)\to\{-1,1\}\) is \(k\)-local positive if every connected subgraph with exactly \(k\) edges has strictly positive total weight, and \(k\) may be forcing, weakly forcing, or collapsing for \(\mathcal G\). For families between trees and all connected graphs, the classification is: \(k=1,2,4\) forcing; \(k=3,5,6,8\) weakly forcing; and \(k=7\) together with all \(k\ge 9\) collapsing [1711.09422]. Although this is not the same process as graph zero forcing, it illustrates the breadth of “forcing” terminology within graph theory.

## 4. \( \mathcal K \)-forcing in Fraïssé-theoretic forcing

In model theory and set theory, Golshani introduces a forcing notion associated with a Fraïssé class \(\mathcal K\) of finite structures in a fixed finite relational language \(\mathcal L\). For an infinite cardinal \(\kappa\),
\[
\mathbb P_{\mathcal K,\kappa}
=
\bigl\{\,p\in\mathcal K : A_p\subseteq \kappa,\ |A_p|<\omega\,\bigr\},
\]
ordered by
\[
p\le q
\iff
\bigl(A_p\supseteq A_q\ \wedge\ p|_{A_q}=q\bigr).
\]
Equivalently, \(p\) extends \(q\) when \(p\) is a strong extension of \(q\) that is the identity on \(A_q\) [1902.06206].

The forcing satisfies the countable chain condition, and in fact is \(\omega_1\)-Knaster if \(\kappa\ge \omega_1\). The stated proof uses the \(\Delta\)-system lemma together with the amalgamation property of \(\mathcal K\). A corollary is that forcing with \(\mathbb P_{\mathcal K,\kappa}\) preserves all cardinals and cofinalities [1902.06206].

If \(G\subseteq \mathbb P_{\mathcal K,\kappa}\) is \(V\)-generic, one defines
\[
K^G=\bigcup_{p\in G} A_p\subseteq \kappa
\]
and, for each relation symbol \(R\in\mathcal L\),
\[
R^G=\bigcup_{p\in G} R^p.
\]
This gives the structure
\[
M_G=(\kappa,(R^G)_{R\in\mathcal L}).
\]
By density of conditions containing any prescribed \(\alpha\in\kappa\), one gets \(K^G=\kappa\). The generic structure has size \(\kappa\), and every finite member of \(\mathcal K\) embeds into \(M_G\) [1902.06206].

The special case \(\kappa=\omega\) recovers classical Fraïssé theory. Then \(\mathbb P_{\mathcal K,\omega}\) is countable, and one uses the dense sets \(D_n\), \(D_{A,A',f}\), and \(D'_{B,B',f}\) together with the Rasiowa–Sikorski lemma to obtain a generic filter meeting all of them. The resulting \(M_G\) is countable, universal for \(\mathcal K\), and ultrahomogeneous, hence exactly the classical Fraïssé limit \(\mathrm{Flim}(\mathcal K)\) [1902.06206].

The running example is the class of finite linear orders. In that case a condition is a finite linearly ordered set \((A_p,\le_p)\), the forcing remains c.c.c., and in the generic extension \(\le^G=\bigcup_{p\in G}\le_p\) is a \(\kappa\)-dense linear order without endpoints. When \(\kappa=\omega\), the construction yields \((\mathbb Q,\le)\) [1902.06206].

## 5. \(K\)-forcing and strongly \(K\)-forcing in \((0,1)\)-matrix theory

For a fixed \(s\times t\) \((0,1)\)-matrix \(K\), an \(m\times n\) matrix \(A\) with \(m\ge s\) and \(n\ge t\) is \(K\)-forcing if every choice of \(s\) rows and \(t\) columns of \(A\) produces an \(s\times t\) submatrix \(A'\) from which one can turn some \(1\)'s to \(0\)'s and obtain exactly \(K\). Equivalently, every \(s\times t\) submatrix of \(A\) covers \(K\) in the usual pattern-containment sense. The extremal function
\[
m(m,n,K)
\]
is the minimum number of \(1\)-entries in an \(m\times n\) \(K\)-forcing matrix [2510.27076].

The paper proves existence and uniqueness of a minimizer. There is a unique \(m\times n\) \(K\)-forcing matrix \(A_{\min}\) with the fewest \(1\)'s, obtained by starting from the all-zero matrix and, for each placement of an \(s\times t\) window in the \(m\times n\) grid, forcing to \(1\) every position of that window whose corresponding entry in \(K\) is \(1\). From this construction one gets monotonicity in the pattern:
if \(K_1\) and \(K_2\) are the same size and \(K_1\le K_2\) entrywise, then
\[
m(m,n,K_1)\le m(m,n,K_2)
\]
[2510.27076].

A geometric description is given in terms of the positions that remain \(0\). The relevant objects are the four corner-functions \(NW(K)\), \(NE(K)\), \(SW(K)\), and \(SE(K)\), each defined as a largest set of zero-positions not dominating any \(1\)-entry in the appropriate corner orientation. These are described as Young diagrams of zeros “scooped out” at a corner of \(K\). When \(m\ge 2s\), \(n\ge 2t\), and \(K\) has no all-zero boundary row or column,
\[
m(m,n,K)=mn-\bigl(|NW(K)|+|NE(K)|+|SW(K)|+|SE(K)|\bigr).
\]
In the general rectangular case one must also subtract linear corrections in \(m\) and \(n\) coming from all-zero boundary runs of \(K\). The worked example is a \(7\times 6\) pattern with corner sizes \(7,4,8,2\), for which the minimizer has
\[
m(m,n,K)=mn-21
\]
for every \(m\ge 14\), \(n\ge 12\), once the boundary-zero runs are accounted for [2510.27076].

The paper also defines strongly \(K\)-forcing. An \(m\times n\) matrix \(A\) is strongly \(K\)-forcing if every \(1\)-entry of \(A\) lies inside some \(s\times t\) submatrix of \(A\) that is exactly equal to \(K\). The associated extremal function
\[
M(m,n,K)
\]
is the maximum number of \(1\)-entries in such a matrix. Unlike the \(K\)-forcing case, exact monotonicity in \(m\) or \(n\) fails, but there is a universal linear-deficit estimate:
\[
mn-M(m,n,K)=O(m+n).
\]
Thus the minimum possible number of \(0\)-entries in a strongly \(K\)-forcing matrix is always linear in the side lengths [2510.27076].

Exact formulas are proved for several permutation patterns. For the \(2\times 2\) identity and anti-identity,
\[
M(n,I_2)=n^2-n,
\qquad
M(n,H_2)=n^2-n,
\]
with unique extremal matrices \(J_n-H_n\) and \(J_n-I_n\), respectively. For every \(3\times 3\) permutation matrix \(P\) and all \(n\ge 3\),
\[
M(n,P)=n^2-3n+3.
\]
For larger identities \(I_k\), the paper gives the lower-bound construction
\[
S_{n,k}=I_{k-2}\oplus (J_{n-k+2}-H_{n-k+2}),
\]
with
\[
|S_{n,k}|=n^2-(2k-3)n-(2k-k^2),
\]
and conjectures that, for every \(k\ge 2\) and \(n\ge k\),
\[
M(n,I_k)=n^2-(2k-3)\,n-(2k-k^2)
\]
[2510.27076].

## 6. K-Forcing in push-forward language modeling

In efficient language generation, K-Forcing is introduced as a joint next-\(k\)-token decoding paradigm for autoregressive language models. Standard autoregressive sampling factorizes
\[
p(x_{t+1:t+k}\mid x_{1:t})
=
\prod_{j=1}^k p_{\rm AR}(x_{t+j}\mid x_{1:t+j-1}),
\]
so generating \(k\) tokens requires \(k\) forward passes. K-Forcing instead distills the autoregressive teacher into a push-forward language model \(G_\theta\) that takes the same context and an independent uniform noise vector
\[
\mathbf z=(z_1,\dots,z_k)\sim \mathrm{Uniform}([0,1]^k)
\]
and returns a joint sample
\[
G_\theta(x_{1:t},\mathbf z)=
(\hat x_{t+1},\dots,\hat x_{t+k})\in \mathcal V^k.
\]
The paper describes this as trading \(k\) memory-bound autoregressive evaluations for a single evaluation that emits a fixed block of \(k\) tokens [2606.10820].

The idealized target is an inverse-CDF push-forward sampler:
\[
\hat x_{t+j}
=
F_{\rm AR}^{-1}\!\bigl(z_j\mid x_{1:t},\hat x_{t+1:t+j-1}\bigr),
\qquad j=1,\dots,k.
\]
Since implementing this exactly would still require \(k\) autoregressive calls, the student model is trained to approximate it. With output heads \(p_{\theta,j}(\cdot\mid x_{1:t},\mathbf z)\), the objective is the next-\(k\)-token prediction cross-entropy
\[
\mathcal L_{\rm PFLM}(\theta)
=
-\frac{1}{|\mathcal T|\,k}
\sum_{t\in\mathcal T}\sum_{j=1}^k
\log p_{\theta,j}\bigl(\hat x_{t+j}\mid x_{1:t},\mathbf z^{(t)}\bigr).
\]
Teacher-generated targets are obtained using sampled noise vectors \(\mathbf z^{(t)}\) [2606.10820].

Training uses progressive self-forcing distillation. Stage 1 distills the autoregressive teacher into a \(k=1\) push-forward model. Stage 2 doubles the window: a teacher PFLM(\(k\)) rolls out twice to produce a \(2k\)-token target block, and the student PFLM(\(2k\)) is trained in one pass to match it. Repeating the doubling schedule \(1\to 2\to 4\to \cdots\) scales the method to larger \(k\) while keeping supervision noise-conditioned [2606.10820].

The implementation uses a fully causal layout in which each noise variable \(z_j\) is embedded as a separate token attending only to the prefix and earlier noise tokens, together with a single shared prediction head. At inference time the system appends \(k\) new uniform noise tokens, performs one forward pass, appends the output tokens’ KV states to the cache, and discards the noise KV. The paper emphasizes that this fixed-stride behavior keeps batch indices and attention masks synchronized and avoids the ragged-tensor problem. A current limitation is training cost: each distillation iteration requires two teacher forward calls plus one student pass, and the present FlashAttention-based implementation uses dense masks with \(O(k^2)\) time instead of the \(O(k)\) block-sparse cost suggested by the mask structure [2606.10820].

Experiments are reported on LM1B and OpenWebText using a 12-layer, approximately \(100\)M-parameter Transformer. Progressive distillation is run in three stages, AR\(\to k=1\), \(1\to 2\), and \(2\to 4\), each for \(500\)K steps. Under bf16 on an NVIDIA H100 at batch sizes \(4\), \(16\), and \(128\), K-Forcing with \(k=4\) gives LM1B throughputs of \(1.69\), \(6.77\), and \(46.5\) k/s, compared with \(0.54\), \(2.03\), and \(15.4\) k/s for the autoregressive baseline. On OpenWebText, the corresponding figures are \(1.70\), \(6.91\), and \(22.9\) k/s against \(0.53\), \(1.99\), and \(9.61\) k/s. The abstract summarizes the aggressive \(k=4\) setting as delivering approximately \(2.4\)-\(3.5\times\) speedup across different batch sizes with modest quality degradation relative to the autoregressive teacher [2606.10820].

The paper also positions K-Forcing against MDLM, Medusa, and PTP draft heads using a quality–NFE comparison on OpenWebText. Its stated interpretation is that K-Forcing achieves the most favorable quality–NFE frontier by modeling joint multi-token blocks rather than independent marginals or draft-and-verify methods whose NFEs double per iteration. The listed strengths are batch-serving compatibility, joint sampling, and a tunable speed–quality trade-off; the listed limitations are a residual quality gap, training overhead, and numerical reproducibility challenges [2606.10820].

Source: https://www.emergentmind.com/topics/k-forcing