---
title: Bottleneck Patching Strategies
url: https://www.emergentmind.com/topics/bottleneck-patching
type: topic
---

# Bottleneck Patching Strategies

Bottleneck patching is a cross-domain label for interventions that do not remove a limiting constraint directly, but alter the states, flows, or computations immediately around that constraint so that system-level degradation is reduced. In the cited literature, the term is used for upstream velocity control before a periodically closing exit in exclusion-process traffic models, assist-warps that repurpose idle GPU resources against a saturated subsystem, transient scratchpads that refine patch-level context in byte-level language models without enlarging persistent KV state, sender-side modifications to BBR under CPU contention, prompt-level repairs of diagnosed modules in multi-stage LLM agents, optimal patch dissemination in clustered malware epidemics, and augmenting-path rewiring in bottleneck assignment [1903.11319; 1602.01348; 2605.09630; 2601.05665; 2605.21958; 1403.1639; 2011.09606]. This breadth suggests a recurring design principle rather than a single algorithm: identify the effective bottleneck state, intervene at the interface that governs it, and optimize the intervention against the induced secondary effects.

## 1. Cross-domain semantics and recurrent structure

Across these works, the bottleneck itself is represented in markedly different mathematical objects. In traffic flow, it is a time-periodic boundary capacity constraint at the exit, with \(\beta(t)\in\{0,1\}\) switching between open and closed phases. In GPU architecture, it is a saturated subsystem such as off-chip memory bandwidth or compute pipelines. In patch-based byte models, it is the coarse patch state that refreshes too slowly relative to the underlying byte stream. In BBR, it is the sender’s CPU time rather than the path bandwidth. In modular LLM agents, it is the module with the largest causal contribution to system failure. In clustered epidemics, it is the type-dependent transmission structure; in bottleneck assignment, it is the maximum-weight edge in a maximum-cardinality matching [1903.11319; 1602.01348; 2605.09630; 2601.05665; 2605.21958; 1403.1639; 2011.09606].

The corresponding patch acts on equally different control variables. The traffic model varies the hopping probability \(v\) through the control parameter \(p\). CABA deploys assist warps through the Assist Warp Store, Assist Warp Controller, and Assist Warp Buffer. Scratchpad Patching introduces transient states \(z_\ell^t\) triggered by entropy. The BBR patch conditionally overwrites \(\text{pacing\_gain}\) with the existing high-gain constant. Counterfactual Correction Patching appends few-shot “Wrong/Correct” demonstrations to one module’s prompt. Epidemic control varies \(u_i(t)\), the patch transmission intensity for each type. The distributed BAP algorithm updates \(\mathcal{M}\) by symmetric difference with an augmenting path, \(\mathcal{M}\oplus\mathcal{P}\) [1903.11319; 1602.01348; 2605.09630; 2601.05665; 2605.21958; 1403.1639; 2011.09606].

A further commonality is that the intervention is almost never evaluated only locally. The relevant objective is global throughput \(Q\), end-to-end performance and energy, downstream evaluation quality, aggregate failure index \(F(E)\), aggregate epidemic cost, or the maximal cost in a complete assignment. This suggests that bottleneck patching is fundamentally about indirect control: local intervention is judged by a nonlocal objective.

## 2. Flow regulation in traffic and transport protocols

In the controlled TASEP with a slow-to-start rule, the bottleneck is encoded by a periodic right-boundary extraction probability
\[
\beta(t)=
\begin{cases}
1 & (nT \le t < nT+\tau),\\
0 & (nT+\tau \le t < (n+1)T),
\end{cases}
\]
with time-averaged exit probability \(\beta^*=\tau/T\). The upstream control is the state-dependent hopping rule
\[
v(t)=
\begin{cases}
0 & \text{if right neighbor is occupied},\\
p & \text{if right neighbor empty and } \beta(t)=0,\\
1 & \text{if right neighbor empty and } \beta(t)=1,
\end{cases}
\]
combined with a slow-to-start coefficient \(s\in[0,1]\) that reduces restart probability after blocking. The mechanism is explicit: when the bottleneck is closed, upstream particles decelerate, which stretches spacing and reduces the discharge penalty induced by slow-to-start when the exit reopens [1903.11319].

The paper distinguishes low-density and high-density regimes. In the high-density regime with \(s=0\), the uncontrolled discharge is strongly reduced: at the exit, one particle leaves every \(3\) steps when congested, and for opening duration \(\tau\) the discharged count per cycle is \(\left\lceil \tau/3 \right\rceil\), yielding
\[
Q=\frac{\left\lceil \tau/3 \right\rceil}{T}.
\]
With control, the theoretical high-density bounds are \(Q_{\min}=\beta^*/3\) and \(Q_{\max}=\beta^*/2\), so the improvement factor can in principle reach \(3/2\), i.e. \(a_{\max}=0.5\). For \(\beta^*=0.6\), \(T=20\), and \(s=0\), simulations reach \(a(p)\approx 0.1\) at an optimal \(p\). The approximate design rules are
\[
p_{\rm opt}=\min\left\{\frac{3}{(1-\beta^*)T},\,1\right\}
\quad\text{and}\quad
l_{\rm opt}\approx \beta^*T,
\]
with a realistic estimate that a VSL-type control based on the scheme could improve intersection throughput by about \(4\%\) [1903.11319].

An analogous logic appears in 2BRobust, but the bottleneck is not a road exit; it is sender CPU time inside a VM. Standard BBR maintains
\[
\text{cwnd}=\text{BtlBw}\cdot \text{RTprop}\cdot \text{cwnd\_gain},
\qquad
\text{pacing\_rate}=\text{BtlBw}\cdot \text{pacing\_gain}.
\]
Under CPU contention, the VM is frequently descheduled, BBR misses pacing opportunities, inflight stays below \(1\) BDP, the measured delivery rate drops, and the controller enters a downward spiral in which \(\text{BtlBw}\), \(\text{pacing\_rate}\), and throughput collapse. The measurements show that CPU-limited BBR senders are capped at very low throughput levels below \(10\)–\(20\) Mbps under any hypervisor and all tested BDP conditions, whereas Cubic remains robust [2601.05665].

The patch is intentionally minimal. Let
\[
\text{BDP}_{\text{est}}=\text{BtlBw}\cdot \text{RTprop}.
\]
If the socket is not app-limited and
\[
\text{inflight}<\text{BDP}_{\text{est}},
\]
the implementation overrides the current phase gain by setting \(\text{pacing\_gain}=\text{HIGH\_GAIN}\) inside `bbr_update_gains()`. This raises both pacing rate and TSO burst size during on-CPU windows. At \(100\) Mbps, with \(1\) ms timeslice, original BBRv3 needs about \(45\%\) CPU share before median throughput reaches at least \(80\) Mbps, while patched BBRv3 recovers near line rate already at about \(25\%\) CPU share; with \(10\) ms timeslice the corresponding threshold shifts from about \(70\%\) to about \(40\%\). The patch does not solve the most severe cases of scheduling, but it solves the problem for the most critical cases and is inert when no inflight deficit occurs [2601.05665].

These two cases instantiate the same systems-level idea at different abstraction levels. In both, the patch does not enlarge the bottleneck directly; it reshapes arrivals into it.

## 3. Architectural bottleneck patching in GPUs

Core-Assisted Bottleneck Acceleration treats bottleneck patching as dynamic conversion of slack in one GPU subsystem into relief for another. The motivating observation is empirical: for \(17\) out of \(27\) applications, the GPU is memory-bound, and at baseline bandwidth memory plus data-dependence stalls constitute about \(61\%\) of issue cycles; doubling bandwidth reduces that to about \(51\%\). At the same time, occupancy constraints leave substantial on-chip state unused; for a \(128\) KB register file, around \(24\%\) of registers are statically unallocated on average. CABA exploits precisely this mismatch between the saturated subsystem and idle resources [1602.01348].

The mechanism is the assist warp. An assist warp shares warp ID and context with its parent warp, executes code sequences stored in the Assist Warp Store, is managed by the Assist Warp Controller through entries in the Assist Warp Table, and is staged in the Assist Warp Buffer. High-priority assist warps, such as decompression on a compressed load return, can block the parent until completion; low-priority assist warps, such as compression or memoization maintenance, are scheduled only in otherwise idle cycles. The AWC also monitors pipeline utilization so that assist deployment is throttled when it would compete with the currently saturated resource [1602.01348].

The most detailed instantiation is bandwidth compression. Between L2 and DRAM, data are stored in compressed form, while L1 is left uncompressed by default. For BDI, a \(64\) B line can be compressed to \(17\) B, yielding
\[
CR=\frac{64}{17}\approx 3.76.
\]
Decompression is implemented as a vector add of base and deltas by a high-priority assist warp; compression is off the critical path and is handled by low-priority assist warps. CABA also adapts FPC and C-Pack, despite their variable-length encodings, by restructuring metadata and enforcing more SIMD-friendly layouts [1602.01348].

The reported performance figures are specific. Across memory-bandwidth-sensitive applications, CABA-BDI improves performance by \(41.7\%\) on average, up to \(2.6\times\); it is only \(2.8\%\) slower than Ideal-BDI and \(1.6\%\) slower than HW-BDI, while outperforming HW-BDI-Mem by about \(9.9\%\). Average DRAM bus utilization is reduced from about \(53.6\%\) to about \(35.6\%\). Total system energy is reduced by about \(22.2\%\), including about \(29.5\%\) average DRAM power reduction, and EDP is reduced by \(45\%\), even though average power increases by about \(2.9\%\). The framework is also described as applicable to memoization, prefetching, redundant multithreading, speculative precomputation, profiling, and reliability tasks, provided that the helper work is mapped onto non-bottleneck resources [1602.01348].

CABA therefore exemplifies bottleneck patching as microarchitectural work migration: compute is spent to simulate bandwidth, or memory activity is spent to relieve compute.

## 4. Representational bottlenecks in patch-based language models

Scratchpad Patching addresses a different bottleneck: the coarse update schedule induced by patch-based byte modeling. A byte-level autoregressive LM is factored into an encoder \(\mathcal{E}\), a patchifier \(\mathcal{P}\), a main trunk \(\mathcal{M}\) running on patch states, an unpatchifier \(\mathcal{U}\), and a decoder \(\mathcal{D}\). In the baseline unpatchifier, for byte position \(n\) in patch \(\ell\),
\[
u_n=
\begin{cases}
\tilde z_{\ell-1}+x_n, & n\neq e_\ell,\\
\tilde z_\ell+x_n, & n=e_\ell.
\end{cases}
\]
Hence every nonterminal byte in a patch must use stale context from the previous patch. The paper identifies this as patch lag: as average patch size grows, compute and KV cache decrease, but modeling quality degrades because patch-level context is refreshed too infrequently [2605.09630].

Scratchpad Patching inserts transient scratchpads inside each patch. With binary trigger \(p_n\in\{0,1\}\), the number of scratchpads in patch \(\ell\) is
\[
T_\ell=\sum_{j=s_\ell}^{e_\ell} p_j,
\]
and the \(t\)-th scratchpad state is
\[
z_\ell^t=\operatorname{Aggregate}(x_{s_\ell:n_t}).
\]
During training, the trunk sees
\[
[z_0,\; z_1^1,\dots,z_1^{T_1},z_1,\; z_2^1,\dots,z_2^{T_2},z_2,\dots],
\]
under a specialized mask: states belonging to patch \(\ell\) may attend to themselves and committed states of earlier patches only; scratchpads neither attend to each other nor are attended to by other elements. During inference, scratchpads are computed on the fly and immediately discarded; only the final committed patch state remains in the KV cache [2605.09630].

The trigger policy most emphasized in the paper is entropy-based. An auxiliary LM head predicts next-byte entropy
\[
H_n=-\sum_{b\in\mathcal{V}} p_\theta(b\mid x_{\le n})\log p_\theta(b\mid x_{\le n}),
\]
and a scratchpad fires when
\[
p_n=\mathbf{1}_{[H_n>\tau_{\text{SP}}]}.
\]
This makes scratchpad density an inference-time compute knob: lower \(\tau_{\text{SP}}\) yields more transient refinement and less patch lag; higher \(\tau_{\text{SP}}\) yields fewer scratchpads and lower compute [2605.09630].

The key claim is decoupling. Patch size still controls the number of committed states \(L\) and thus the KV-cache footprint, but scratchpad density controls additional compute without increasing persistent memory. The reported results are concrete. At \(16\) bytes per patch, SP-augmented models match or closely approach the byte-level baseline on downstream evaluations while using a \(16\times\) smaller KV cache over patches and \(3\)–\(4\times\) less inference compute. On natural-language NLU, fixed-size \(p=16\) without SP yields average score about \(48.0\), whereas fixed \(p=16\) with SP yields about \(54.2\), matching the byte-level baseline at \(54.1\). On MBPP pass@1, byte-level reaches \(26.3\), fixed \(p=16\) reaches \(18.2\), and fixed \(p=16\) with SP reaches \(27.5\). The paper also reports that H-Net + SP can trigger redundant compute because scratchpads often fire right before learned boundaries, and that very dense scratchpads give diminishing returns and can even slightly hurt NL BPB [2605.09630].

Within this literature, SP is a particularly explicit formulation of bottleneck patching: the bottleneck state is not removed but supplemented by transient, nonpersistent refinements.

## 5. Diagnostic bottlenecks and patching hazards in modular LLM pipelines

In multi-module LLM agents, bottleneck patching is not presented as a throughput optimization but as a prescription problem after causal diagnosis. The studied agent family has four modules,
\[
M_1 \to M_2 \to M_3 \to M_4,
\]
corresponding to Query Rewriter, Planner, Router, and Response Generator. System performance is summarized by the failure index
\[
F(E)=-\sum_{i=1}^{k}\log\bigl(1-\mathrm{sev}_i\bigr),
\]
and the causal responsibility of module \(M_i\) is
\[
\Delta F_i(c)=F(E(c))-F(E(c)\mid \mathrm{do}(M_i=S_i^*)).
\]
On \(\tau\)-bench retail with the gpt-4o-mini agent, the average effects are \(\overline{\Delta F_1}=0.357\), \(\overline{\Delta F_2}=0.817\), \(\overline{\Delta F_3}=1.018\), and \(\overline{\Delta F_4}=0.589\), so the router \(M_3\) is the diagnosed primary bottleneck; the same qualitative pattern holds across retail, airline, Llama 4 Scout, and Qwen3-32b [2605.21958].

The central result is that diagnosis does not identify the best patch location. Counterfactual Correction Patching constructs a correction pool from diagnosis data, filters examples with \(\mathrm{sev}_i^{(0)}(c)\ge 0.30\), keeps the top \(k=5\), and appends them to one module’s prompt in the form
```text
### Example j
Input: <x_i>
Wrong: <M_i^{(0)}(c)>
Correct: <S_i^*(c)>
```
while leaving instructions and output schema unchanged. Applied to the diagnosed bottleneck \(M_3\), this consistently fails as a prescription. On retail, Pop CCP at \(M_3\) changes mean failure index by \(+0.243\) for gpt-4o-mini, with Holm \(p<0.05\); by \(+0.683\) for Qwen3-32b, with raw \(p<0.001\); and by \(-0.013\) for Llama 4 Scout, i.e. neutral. By contrast, Pop CCP at the upstream Query Rewriter \(M_1\) changes mean failure index by \(-0.191\), \(-0.659\), and \(-0.402\), respectively. Tool-match shows the same asymmetry: for gpt-4o-mini, baseline is \(37.8\%\), CCP at \(M_1\) gives \(44.1\%\), and CCP at \(M_3\) gives \(33.3\%\); for Qwen3-32b, baseline is \(44.9\%\), CCP at \(M_1\) stays at \(44.9\%\), and CCP at \(M_3\) drops to \(24.5\%\) [2605.21958].

The paper explains this through the Linguistic Contract hypothesis. Downstream modules adapt to the characteristic error distribution of upstream modules, so correcting the bottleneck may break that co-adaptation. The mediation quantity is the Natural Indirect Effect,
\[
\mathrm{NIE}_i=F(\text{A})-F(\text{B}),
\]
where World A re-executes \(M_i\) on oracle upstream and World B freezes \(M_i\) to its original output while upstream is corrected. Tasks are then classified as amplifier, propagator, or compensator according to the sign and magnitude of \(\mathrm{NIE}_i\). For gpt-4o-mini retail, \(M_3\) has \(\overline{\mathrm{NIE}_3}=-1.014\) and fate counts \(2/7/491\) for amplifier/propagator/compensator. The per-agent co-adaptation proxy is the compensator rate at \(M_3\): \(98.2\%\) for gpt-4o-mini, \(96.0\%\) for Qwen3-32b, and \(0\%\) for Llama 4 Scout. High co-adaptation co-occurs with CCP hazard; zero co-adaptation co-occurs with neutrality [2605.21958].

The paper further isolates the hazard to correction injection rather than any intervention at the router. Instruction rewriting at \(M_3\) is essentially neutral, with \(\Delta=-0.014\) for gpt-4o-mini; a model upgrade at \(M_3\) is also neutral, with \(\Delta\approx 0\). Oracle injection at \(M_3\), however, strongly improves outcomes, with \(\Delta=-2.007\) for gpt-4o-mini and \(\Delta=-1.795\) for Llama 4 Scout. This indicates that the router is genuinely highly improvable, but that partial prompt-level correction is a hazardous way to intervene there. A further mechanistic observation is that the final-output sentence-embedding cosine is \(0.906\) for both CCP @ \(M_3\) and CCP @ \(M_1\), so the asymmetry is not explained by the magnitude of final distribution shift but by its direction: \(M_3\) perturbs executable-semantic variables such as tool name and argument patterns, whereas \(M_1\) perturbs a surface-linguistic layer that downstream modules tolerate better [2605.21958].

## 6. Optimal-control and combinatorial formulations

The most explicit mathematical theory of patching appears in clustered malware epidemics. The network is partitioned into \(M\) types, with state fractions \(S_i(t)\), \(I_i(t)\), and \(R_i(t)\), infection rates \(\beta_{ji}\), patch dissemination rates \(\bar\beta_{ji}\), healing efficacies \(\pi_{ji}\), and control inputs \(u_j(t)\in[0,1]\). The objective is to minimize aggregate damage and patching cost over \([0,T]\), with cost functionals such as
\[
J_{\text{non-rep}}(\mathbf{u})=
\int_0^T\left(
f(\mathbf{I}(t))-L(\mathbf{R}(t))
+\sum_{i=1}^M R_i^0\,h_i(u_i(t))
\right)\,dt
\]
for non-replicative patching, and the analogous \(R_i(t)\)-weighted form for replicative patching. Applying Pontryagin’s Maximum Principle yields scalar minimization conditions of the form
\[
u_i(t)\in \arg\min_{0\le x\le 1}\psi_i(x,t),
\qquad
\psi_i(x,t)=R_i^0\bigl(h_i(x)-\phi_i(t)x\bigr)
\]
or
\[
\psi_i(x,t)=R_i(t)\bigl(h_i(x)-\phi_i(t)x\bigr),
\]
with \(\phi_i(t)\) strictly decreasing in time. This monotonicity drives the central structural result: if \(h_i\) is concave, the optimal control is bang-bang with at most one switch from \(u_i=1\) to \(u_i=0\); if \(h_i\) is strictly convex, the optimal control is continuous, strictly decreasing from \(1\) to \(0\), with at most one transition interval. Numerical examples on linear, star, and complete topologies show that the optimal stratified dynamic policy can reduce aggregate cost by about \(40\%\) versus the best static stratified policy and about \(100\%\) versus the homogeneous approximation for \(M=5\), while replicative patching can achieve up to about \(60\%\) cost reduction relative to non-replicative patching [1403.1639].

A combinatorial analogue appears in the distributed Bottleneck Assignment Problem. Here the objective is
\[
\min_{\mathcal{M}\in\mathcal{C}(\mathcal{G}_b)}\max_{e\in\mathcal{M}} w(e),
\]
where \(\mathcal{C}(\mathcal{G}_b)\) is the set of maximum-cardinality matchings. The pruneBAP algorithm starts from an MCM \(\mathcal{M}\), selects a current bottleneck edge \(\bar e\in \arg\max_{e\in\mathcal{M}} w(e)\), forms the pruned edge set
\[
\phi(\mathcal{G}_b,\mathcal{M})
=
\mathcal{M}\cup
\{e\in\mathcal{E}_b\mid w(e)<\max_{e'\in\mathcal{M}}w(e')\},
\]
removes \(\bar e\), and searches for an augmenting path relative to \(\mathcal{M}\setminus\{\bar e\}\). If such a path exists, the matching is patched by symmetric difference; if not, \(\bar e\) is critical and the current matching is optimal. The distributed implementation decomposes this into MaxEdge, local pruning, and distributed augmenting-path search via AugDFS or AugBFS over a communication graph of diameter \(D\). For \(m=n\), the worst-case complexity is \(\mathcal{O}(n^3 D)\). AugDFS tends to require fewer pruneBAP iterations, whereas AugBFS tends to require fewer time steps overall; the choice is therefore a communication-versus-time trade-off. The same augmenting-path logic also governs when two independently solved sub-BAPs can be merged directly: if no beneficial cross-cluster augmenting path exists, \(\mathcal{M}_3=\mathcal{M}_1\cup\mathcal{M}_2\) is already a bottleneck assignment for the combined graph [2011.09606].

Taken together, these formulations show that bottleneck patching can be cast either as continuous-time optimal control with monotone shadow prices or as discrete augmenting-path repair on a thresholded graph. A plausible implication is that the most transferable element across domains is not the surface mechanism—speed limit, assist warp, scratchpad, prompt correction, epidemic dispatcher, or matching update—but the underlying strategy of exploiting a local control surface whose marginal value can be tracked against a global bottleneck objective.

Source: https://www.emergentmind.com/topics/bottleneck-patching