---
title: Multi-Level Transversal Injection (MLTI)
url: https://www.emergentmind.com/topics/multi-level-transversal-injection-mlti
type: topic
---

# Multi-Level Transversal Injection (MLTI)

Multi-Level Transversal Injection (MLTI) most directly denotes a fault-tolerant protocol for preparing logical rotation states at arbitrary Clifford hierarchy levels by iterating logical-level transversal injection and magic-state pumping [2509.23642]. The same label, or closely related inferred abstractions, also appears in surface-code ancilla preparation, one-way-transversal code switching, transversal STAR gadgets, transformer depth upscaling, transversal Hamiltonicity, and stateless multi-turn attack models, where the common motif is a level-wise insertion or injective assignment constrained by a transversal structure [2404.01301][2409.13465][2509.18294][2410.11654][2111.07079][2604.21860]. *This suggests* that MLTI is not a single domain-independent construction, but a family of techniques built around structured injection across multiple layers, levels, or subsystems.

## 1. Terminology and scope

The expression “MLTI” is explicit in the 2025 letter on logical non-Clifford state preparation, where it names a concrete lattice-surgery-based method for preparing rotation states with overhead that decreases with Clifford hierarchy level and then plateaus [2509.23642]. In several other works, by contrast, the term is not part of the paper’s formal title or original notation; instead, it is introduced as an inferred abstraction consistent with the paper’s formalism, such as hierarchical transversal state preparation in stabilizer codes, multi-stage transversal STAR gadgets, or one-way-transversal code-switching pipelines [2404.01301][2211.10046][2509.18294][2409.13465].

| Context | Meaning of injection | Representative paper |
|---|---|---|
| Fault-tolerant quantum computing | Logical-level preparation or transfer of non-Clifford resources across code levels or patches | [2509.23642] |
| Extremal combinatorics | Injection $\varphi$ assigning edges of a spanning structure to distinct layers | [2111.07079] |
| Transformer scaling | Regular insertion of new layers across depth with near-identity initialization | [2410.11654] |
| LLM safety | Distribution of adversarial intent across turns or system levels | [2604.21860] |

Within quantum computing, “transversal” retains its standard FT connotation: operations are arranged so that error propagation is constrained by code structure, while the injected object is typically a resource state, a non-Clifford phase, or a temporary encoding transfer. Within combinatorics, “transversal” refers to selecting one edge from each layer, or more precisely assigning distinct edges to distinct layers via an injection. In transformer scaling, the term denotes layer insertion across the stack rather than end-appending. In LLM safety, it denotes adversarial traversal across levels of a pipeline while each local safeguard still returns `allow`.

## 2. Explicit MLTI for logical rotation-state preparation

In its explicit quantum-information sense, MLTI is a protocol for preparing rotation states at arbitrary Clifford hierarchy levels. The basic single-qubit objects are the $Z$-axis rotation
$$
U_Z(\theta)=R_z(\theta)=e^{i\theta Z}
$$
and the corresponding rotation state
$$
|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.
$$
A level-$l$ rotation is $R_z(\pi/2^l)$, and the associated state $|\pi/2^l\rangle$ is a level-$l$ rotation state [2509.23642].

The core physical-level primitive is transversal injection of the form $k|\alpha\rangle \to |\beta\rangle$. Starting from $|+_L\rangle$, one applies $R_z(\alpha)$ to each of the $k$ physical qubits in $Z_L$. In the noiseless case, post-selection yields
$$
|\beta_L\rangle
=
\frac{\cos^k \alpha \, |+_L\rangle + i \sin^k \alpha \, |- _L\rangle}{p_s^{(0)}},
$$
with
$$
\beta = \arctan(\tan^k \alpha) \approx \alpha^k,
\qquad
p_s^{(0)}=\cos^{2k}\alpha+\sin^{2k}\alpha.
$$
Under circuit-level noise with physical error rate $p_{\mathrm{phy}}$, the output infidelity scales as
$$
\epsilon_L \propto k \beta^{2(1-1/k)} p_{\mathrm{phy}} + O(p_{\mathrm{phy}}^2).
$$
MLTI lifts this primitive to the logical level by using lattice surgery to measure $X_iX_{i+1}$ across $k$ logical patches, thereby implementing the same ideal map $k|\alpha\rangle \to |\beta\rangle$ on logical inputs [2509.23642].

The central suppression statement is the per-level theorem:
$$
\epsilon_L \le k^2 \beta^{2(1-1/k)} \epsilon + O(\epsilon^{1.5}),
$$
where $\epsilon$ is the input infidelity to $|\alpha\rangle$. This is the mechanism by which successive logical levels reduce error. Because repeated injection would otherwise shrink the angle too aggressively, MLTI introduces magic-state pumping: after each level, apply $R_z(-\pi/8)$ and $X$ so that the next level’s input angle is brought close to $\pi/8$ in magnitude. The protocol therefore alternates suppression and angle restoration [2509.23642].

The implementation is surface-code based and uses asymmetric rotated patches with $d_z>d_x$, since a $Z$ error sends $|\theta\rangle$ to $|\theta^\perp\rangle$ and therefore dominates infidelity, whereas $X$ errors contribute only $\sin^2 2\theta$. The logical error estimates used in the resource model are
$$
P_L(d)\approx 0.1 (100 p_{\mathrm{phy}})^{(d+1)/2},
$$
with
$$
P_L^x(d,d)=P_L^z(d,d)\approx 0.05 (100 p_{\mathrm{phy}})^{(d+1)/2},
$$
and for rectangular patches the paper rescales by area ratios:
$$
P_L^x(d_x,d_z)\approx \frac{d_x d_z}{d_x^2} P_L^x(d_x,d_x),
\qquad
P_L^z(d_x,d_z)\approx \frac{d_x d_z}{d_z^2} P_L^z(d_z,d_z).
$$
This asymmetry is part of the overhead reduction, not a secondary optimization [2509.23642].

A second technical contribution is elimination of off-diagonal terms “for free.” Naïve dephasing in the $\{|\theta\rangle,|\theta^\perp\rangle\}$ basis would require an extra $R_z(2\theta)$, but MLTI integrates dephasing into teleportation by randomly pre-applying $X$ to the ancilla with probability $1/2$ and switching between two teleportation channels, $\mathcal{G}$ and $\mathcal{G}'$. This cancels off-diagonal contributions without introducing extra rotations, thereby making infidelity and trace distance coincide for the prepared ancilla [2509.23642].

The resource model is stated in space-time volume, defined as physical-qubit count times QEC cycles per successfully produced target state. For the first level,
$$
V^{(1)}=(2 d_z^{(1)} d_x^{(1)}) \times 2 / p_{s,1}.
$$
For pumping,
$$
V^{(1)}_{\mathrm{pump}}
=
V_T^{(1)} + 2 d_z^{(1)}(d_x^{(1)}+d_z^{(1)}+1)\times (2d_z^{(1)}+\lfloor d_z^{(1)}/2\rfloor+2).
$$
For the second level,
$$
V^{(2)}
=
\bigl[2 d_x^{(2)} (k^{(2)} d_z^{(2)} + k^{(2)} - 1)\times(d_z^{(2)}+2) + V^{(1)} + V^{(1)}_{\mathrm{pump}}\bigr]/p_{s,2}.
$$
Quantitatively, at $p_{\mathrm{phy}}=5\times 10^{-4}$ and target infidelity $<10^{-12}$, the reported space-time volume is below state distillation for all $l>7$, and for $l>12$ it is lower by a factor $>10$; with up to MLTI level $r=4$, high-fidelity preparation of all rotation states with $l>6$ is achievable [2509.23642].

## 3. Quantum antecedents and related architectures

MLTI in the 2025 sense sits within a broader line of transversal-injection ideas. In surface-code state preparation, Transversal Injection (TI) initializes every data qubit in
$$
|\chi\rangle = \alpha |0\rangle + \beta |1\rangle
$$
before standard stabilizer measurements. The state $|\chi\rangle^{\otimes N}$ is projected by the stabilizer trajectory into a logical non-Pauli state, and the logical amplitudes are computed by trajectory-dependent sums of monomials $\alpha^j \beta^{N-j}$ [2404.01301]. A related stabilizer-code treatment states the same mechanism in terms of uniform physical rotations, trajectory-conditioned amplitude sums, and gate teleportation of $R_z(\theta)$ or $R_x(\theta)$ resource states; its hierarchical extension is described there as a multi-level application of TI across concatenated levels [2211.10046]. In both accounts, the preparation is probabilistic and heralded, and the logical output depends on measured syndrome history rather than being fixed a priori.

Other FT constructions instantiate the same multi-stage pattern without using the term explicitly. One-way-transversal code switching realizes a logical $T$ gate by teleporting from a self-dual code to a triorthogonal code, applying a transversal $T$, and teleporting back using only transversal CNOTs, transversal single-qubit rotations, and transversal measurements [2409.13465]. For the $d=3$ pair, the protocol uses the Steane code $[[7,1,3]]$ and Tetrahedral code $[[15,1,3]]$; its logical failure rate scales as $L(p)=O(p^2)$, with break-even against a physical $T$ gate at approximately $p\approx 2\times 10^{-3}$, while the $d=5$ variants achieve $L(p)=O(p^3)$ [2409.13465]. *This suggests* an MLTI interpretation in which transversal transfer between code levels substitutes for a dedicated magic-state factory.

A second related architecture is transversal STAR for neutral-atom simulation. There the inferred “multi-level” structure consists of four stages: transversal multi-rotation injection inside a patch, repeat-until-success teleportation to data patches, composition under correlated decoding with $O(1)$ syndrome extraction between transversal Cliffords, and extension to high-rate codes with fold/swap-transversal Clifford layers [2509.18294]. The small-angle scaling reported for injection is
$$
p_L(\theta)\approx \alpha p_{\mathrm{phys}} |\theta|,
\qquad
\alpha \approx 1.5 \text{ at } d=7,
$$
and the simulation-volume bound is
$$
N_s T \lesssim \frac{1}{l_1 \alpha p_{\mathrm{phys}}}.
$$
At $p_{\mathrm{phys}}=10^{-3}$, the paper reports $N_sT>600$ with approximately $10{,}000$ physical qubits, corresponding to a fully fault-tolerant computation requiring over $10^6$–$10^7$ $T$ gates [2509.18294].

The qutrit literature supplies an additional variant. “Transversal AND in Quantum Codes” constructs a qutrit $[[6,2,2]]$ CSS code with a built-in transversal implementation of AND, derived from a symmetric T-depth-one compute–phase–uncompute circuit, and then concatenates it to a $[[48,2,4]]$ code preserving the same logical gate [2603.04548]. The paper does not name this MLTI, but its layered realization of non-Clifford action through code construction and concatenation is formally close to the quantum uses above.

## 4. Transversal injections in extremal combinatorics

In extremal combinatorics, MLTI denotes injective assignment across layers rather than state preparation. For a $k$-graph system
$$
\mathbf{H}=\{H_i\}_{i\in[m]}
$$
on a common $n$-vertex set $V$, a $k$-graph $H$ is $\mathbf{H}$-transversal if there exists an injection
$$
\varphi:E(H)\hookrightarrow [m]
$$
such that $e\in E(H_{\varphi(e)})$ for all $e\in E(H)$. In this setting, the injection itself is the transversal object [2111.07079].

The main theorem for hypergraph systems states that for $k\ge 3$, $\gamma>0$, sufficiently large $n$, and an $n$-vertex $k$-graph system $\mathbf{H}=\{H_i\}_{i\in[n]}$, if
$$
\delta_{k-1}(H_i)\ge (1/2+\gamma)n
$$
for every $i\in[n]$, then there exists an $\mathbf{H}$-transversal tight Hamilton cycle [2111.07079]. The condition $m=n$ matches the fact that a tight Hamilton cycle has exactly $n$ edges, so an injection $\varphi:E(H)\hookrightarrow [n]$ is exactly what is needed for one edge per layer. The proof uses the absorption method, with a transversal absorbing lemma, a connecting lemma, and a path-cover lemma, together with an auxiliary $(1,k)$-graph $H^*$, weak hypergraph regularity, reduced hypergraphs, and rainbow matching arguments that are then blown up to long transversal paths [2111.07079].

A bipartite analogue appears in collections
$$
\mathbf{G}=\{G_1,\dots,G_{2n-1}\}
$$
on a common bipartition $V=(X,Y)$ with $|X|=|Y|=n$. If $P$ is a Hamiltonian path with $|E(P)|=2n-1$ and there exists an injection
$$
\varphi:E(P)\to [2n-1]
$$
such that each edge lies in its assigned level, then $P$ is a $\mathbf{G}$-transversal isomorphic to a Hamiltonian path [2601.17758]. Theorem 1.3 states that if
$$
\delta(G_i)\ge \left\lceil \frac{n}{2}\right\rceil
$$
for each $i\in[2n-1]$, then either such a transversal Hamiltonian path exists or, when $n$ is even, all levels are the extremal disconnected graph
$$
K_{\frac n2,\frac n2}\cup K_{\frac n2,\frac n2}.
$$
Theorem 1.4 raises the degree threshold to
$$
\delta(G_i)\ge \left\lceil \frac{n+1}{2}\right\rceil
$$
and obtains Hamiltonian connectedness unless the odd-$n$ extremal family $\{F,F'\}$ occurs [2601.17758]. Here MLTI is equivalent to a full rainbow labeling of all $2n-1$ edges of a spanning path.

These combinatorial uses are structurally strict: if the number of layers is below the number of required edges, full transversality is impossible. That exact impossibility statement is explicit for tight Hamilton cycles when $m<n$ and for bipartite Hamiltonian paths when fewer than $2n-1$ layers are available [2111.07079][2601.17758].

## 5. Layer-wise insertion in transformers and multi-level attacks on LLM systems

In transformer scaling, MLTI appears as the general strategy of inserting new transformer layers transversally across the stack at regular intervals, with Transformer Layer Injection (TLI) as the dense-transformer instantiation [2410.11654]. If the base model has depth $L$, injection interval $K$, and injection set
$$
I=\{i\in\{1,2,\dots,L\}\mid i\bmod K=0\},
$$
then the new depth is
$$
L' = L + \lfloor L/K \rfloor.
$$
Each inserted block is initialized to be identity-preserving by duplicating the previous layer and zeroing the attention output projection and FFN down projection, so that the residual branch is initially zero. In pre-norm notation, if
$$
g(h)=h+\alpha\cdot u(h),
$$
TLI uses $\alpha=0$ on the output projections, producing exact identity at initialization [2410.11654].

The training schedule is two-stage: first freeze the original $L$ layers and train only the injected $L/K$ layers, then optionally fine-tune all layers with LoRA or QLoRA. Reported experiments use LLama3 1B, 3B, and 8B on KoBEST and KMCQA. The paper reports lower initialization loss than DUS, fewer training steps to reach target performance, superior downstream accuracy, and several settings in which the injected models perform effectively even without additional training [2410.11654]. Compared with Mixture of Experts, the architecture remains dense and avoids routing overhead; compared with DUS, the benefit is attributed to reduced initialization mismatch and preserved hidden-state continuity.

A different systems interpretation arises in LLM safety. “Transient Turn Injection” (TTI) is defined as a stateless multi-turn attack that distributes adversarial intent across isolated interactions, and the paper explicitly maps MLTI as a generalization across levels of an LLM-enabled pipeline: model front-end, tool calls, retrieval, external agents, content filters, and approval workflows [2604.21860]. In the formalization given there, each level $\ell$ has a component $F^{(\ell)}$ and safeguard $S^{(\ell)}$, and the adversary seeks a final output $y_k^{(L)}\in O$ while all local filters still return `allow`. Under this mapping, TTI is a subclass of MLTI specialized to turns at a single interface.

The evaluation reported for TTI shows substantial variation in robustness across models. In the cited run, Claude 3.5 Haiku has a safe response count of 49/50, GPT-4.1-mini and GPT-4o variants 46/50, while Gemini 2.0 Flash, Gemini 2.0 Flash-Lite, and Gemini 1.5 Flash range from 30/50 to 33/50 safe responses; TTI counts also exceed PAIR counts across models such as Gemini 2.0 Flash (PAIR=4, TTI=34) and GPT-4.1-mini (PAIR=2, TTI=8) [2604.21860]. The mitigations proposed are session-level context aggregation, deep alignment, context-aware moderation, adaptive rate limiting, memory-based consistency checks, and continuous adversarial testing.

## 6. Shared structure, misconceptions, and open directions

A central misconception is that MLTI names a universally standardized object. The record is more heterogeneous. The term is formal and central in the logical rotation-state paper [2509.23642]; it is a named general strategy instantiated by TLI in transformer scaling [2410.11654]; and in several quantum, combinatorial, and safety papers it is an editorial or inferred abstraction consistent with the original formalism rather than the paper’s own headline terminology [2404.01301][2409.13465][2111.07079][2604.21860]. Any domain-independent definition must therefore be treated cautiously.

*This suggests* a shared structural template with three recurring features. First, there is a stratified object: code levels, hypergraph layers, transformer depth positions, or system components. Second, the injection is constrained: one edge per layer, one near-identity block per interval, one logical resource per level, or one locally admissible step per safeguard. Third, performance depends on preserving compatibility with the ambient structure: syndrome-consistent trajectories in FT protocols, codegree or minimum-degree thresholds in Hamiltonicity, hidden-state continuity in transformers, or stateless moderation gaps in LLM pipelines.

The open problems are likewise domain-specific. In explicit quantum MLTI, the paper points to further gains from integrating QLDPC codes and higher-rate architectures, and parameter search over $k^{(r)}$, $d_x^{(r)}$, $d_z^{(r)}$, and pumping fidelities remains architecture-dependent [2509.23642]. In surface-code and code-switching variants, the unresolved bottlenecks include exponential trajectory computation, efficient deterministic auxiliary-state preparation at higher distances, and adaptation to limited-connectivity hardware [2211.10046][2409.13465]. In hypergraph and bipartite transversality, open directions include loose cycles, alternative degree conditions, algorithmic tractability, and extensions to other numbers of levels [2111.07079][2601.17758]. In transformer scaling, practical questions concern faster specialized algorithms, ablations over $K$, and extension beyond depth-only scaling [2410.11654]. In LLM safety, the stated priority is sequence-level and system-level policy reasoning rather than turn-local or component-local filtering, together with continuous automated red-teaming [2604.21860].

Across these literatures, MLTI functions less as a single theorem than as a reusable design principle: inject across levels, preserve a transversal constraint, and exploit the resulting structure to obtain fidelity, Hamiltonicity, scalability, or adversarial evasion, depending on the domain.

Source: https://www.emergentmind.com/topics/multi-level-transversal-injection-mlti