Papers
Topics
Authors
Recent
Search
2000 character limit reached

Multi-Level Transversal Injection (MLTI)

Updated 14 July 2026
  • MLTI is a protocol that prepares logical rotation states across multiple Clifford hierarchy levels using transversal injection and magic-state pumping.
  • It alternates between error suppression and angle restoration to maintain fidelity while reducing physical error contributions.
  • Its design principles extend to diverse fields, including quantum fault tolerance, combinatorial transversality, transformer scaling, and multi-turn LLM safety.

Multi-Level Transversal Injection (MLTI) most directly denotes a fault-tolerant protocol for preparing logical rotation states at arbitrary Clifford hierarchy levels by iterating logical-level transversal injection and magic-state pumping (Zhang et al., 28 Sep 2025). The same label, or closely related inferred abstractions, also appears in surface-code ancilla preparation, one-way-transversal code switching, transversal STAR gadgets, transformer depth upscaling, transversal Hamiltonicity, and stateless multi-turn attack models, where the common motif is a level-wise insertion or injective assignment constrained by a transversal structure (Gavriel et al., 2023, Heußen et al., 2024, Ismail et al., 22 Sep 2025, Vo, 2024, Cheng et al., 2021, Rayhan et al., 23 Apr 2026). This suggests that MLTI is not a single domain-independent construction, but a family of techniques built around structured injection across multiple layers, levels, or subsystems.

1. Terminology and scope

The expression “MLTI” is explicit in the 2025 letter on logical non-Clifford state preparation, where it names a concrete lattice-surgery-based method for preparing rotation states with overhead that decreases with Clifford hierarchy level and then plateaus (Zhang et al., 28 Sep 2025). In several other works, by contrast, the term is not part of the paper’s formal title or original notation; instead, it is introduced as an inferred abstraction consistent with the paper’s formalism, such as hierarchical transversal state preparation in stabilizer codes, multi-stage transversal STAR gadgets, or one-way-transversal code-switching pipelines (Gavriel et al., 2023, Gavriel et al., 2022, Ismail et al., 22 Sep 2025, Heußen et al., 2024).

Context Meaning of injection Representative paper
Fault-tolerant quantum computing Logical-level preparation or transfer of non-Clifford resources across code levels or patches (Zhang et al., 28 Sep 2025)
Extremal combinatorics Injection φ\varphi assigning edges of a spanning structure to distinct layers (Cheng et al., 2021)
Transformer scaling Regular insertion of new layers across depth with near-identity initialization (Vo, 2024)
LLM safety Distribution of adversarial intent across turns or system levels (Rayhan et al., 23 Apr 2026)

Within quantum computing, “transversal” retains its standard FT connotation: operations are arranged so that error propagation is constrained by code structure, while the injected object is typically a resource state, a non-Clifford phase, or a temporary encoding transfer. Within combinatorics, “transversal” refers to selecting one edge from each layer, or more precisely assigning distinct edges to distinct layers via an injection. In transformer scaling, the term denotes layer insertion across the stack rather than end-appending. In LLM safety, it denotes adversarial traversal across levels of a pipeline while each local safeguard still returns allow.

2. Explicit MLTI for logical rotation-state preparation

In its explicit quantum-information sense, MLTI is a protocol for preparing rotation states at arbitrary Clifford hierarchy levels. The basic single-qubit objects are the ZZ-axis rotation

UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}

and the corresponding rotation state

θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.

A level-ll rotation is Rz(π/2l)R_z(\pi/2^l), and the associated state π/2l|\pi/2^l\rangle is a level-ll rotation state (Zhang et al., 28 Sep 2025).

The core physical-level primitive is transversal injection of the form kαβk|\alpha\rangle \to |\beta\rangle. Starting from +L|+_L\rangle, one applies ZZ0 to each of the ZZ1 physical qubits in ZZ2. In the noiseless case, post-selection yields

ZZ3

with

ZZ4

Under circuit-level noise with physical error rate ZZ5, the output infidelity scales as

ZZ6

MLTI lifts this primitive to the logical level by using lattice surgery to measure ZZ7 across ZZ8 logical patches, thereby implementing the same ideal map ZZ9 on logical inputs (Zhang et al., 28 Sep 2025).

The central suppression statement is the per-level theorem:

UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}0

where UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}1 is the input infidelity to UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}2. This is the mechanism by which successive logical levels reduce error. Because repeated injection would otherwise shrink the angle too aggressively, MLTI introduces magic-state pumping: after each level, apply UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}3 and UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}4 so that the next level’s input angle is brought close to UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}5 in magnitude. The protocol therefore alternates suppression and angle restoration (Zhang et al., 28 Sep 2025).

The implementation is surface-code based and uses asymmetric rotated patches with UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}6, since a UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}7 error sends UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}8 to UZ(θ)=Rz(θ)=eiθZU_Z(\theta)=R_z(\theta)=e^{i\theta Z}9 and therefore dominates infidelity, whereas θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.0 errors contribute only θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.1. The logical error estimates used in the resource model are

θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.2

with

θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.3

and for rectangular patches the paper rescales by area ratios:

θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.4

This asymmetry is part of the overhead reduction, not a secondary optimization (Zhang et al., 28 Sep 2025).

A second technical contribution is elimination of off-diagonal terms “for free.” Naïve dephasing in the θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.5 basis would require an extra θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.6, but MLTI integrates dephasing into teleportation by randomly pre-applying θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.7 to the ancilla with probability θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.8 and switching between two teleportation channels, θ=Rz(θ)+=cosθ++isinθ.|\theta\rangle = R_z(\theta)|+\rangle = \cos \theta |+\rangle + i \sin \theta |-\rangle.9 and ll0. This cancels off-diagonal contributions without introducing extra rotations, thereby making infidelity and trace distance coincide for the prepared ancilla (Zhang et al., 28 Sep 2025).

The resource model is stated in space-time volume, defined as physical-qubit count times QEC cycles per successfully produced target state. For the first level,

ll1

For pumping,

ll2

For the second level,

ll3

Quantitatively, at ll4 and target infidelity ll5, the reported space-time volume is below state distillation for all ll6, and for ll7 it is lower by a factor ll8; with up to MLTI level ll9, high-fidelity preparation of all rotation states with Rz(π/2l)R_z(\pi/2^l)0 is achievable (Zhang et al., 28 Sep 2025).

MLTI in the 2025 sense sits within a broader line of transversal-injection ideas. In surface-code state preparation, Transversal Injection (TI) initializes every data qubit in

Rz(π/2l)R_z(\pi/2^l)1

before standard stabilizer measurements. The state Rz(π/2l)R_z(\pi/2^l)2 is projected by the stabilizer trajectory into a logical non-Pauli state, and the logical amplitudes are computed by trajectory-dependent sums of monomials Rz(π/2l)R_z(\pi/2^l)3 (Gavriel et al., 2023). A related stabilizer-code treatment states the same mechanism in terms of uniform physical rotations, trajectory-conditioned amplitude sums, and gate teleportation of Rz(π/2l)R_z(\pi/2^l)4 or Rz(π/2l)R_z(\pi/2^l)5 resource states; its hierarchical extension is described there as a multi-level application of TI across concatenated levels (Gavriel et al., 2022). In both accounts, the preparation is probabilistic and heralded, and the logical output depends on measured syndrome history rather than being fixed a priori.

Other FT constructions instantiate the same multi-stage pattern without using the term explicitly. One-way-transversal code switching realizes a logical Rz(π/2l)R_z(\pi/2^l)6 gate by teleporting from a self-dual code to a triorthogonal code, applying a transversal Rz(π/2l)R_z(\pi/2^l)7, and teleporting back using only transversal CNOTs, transversal single-qubit rotations, and transversal measurements (Heußen et al., 2024). For the Rz(π/2l)R_z(\pi/2^l)8 pair, the protocol uses the Steane code Rz(π/2l)R_z(\pi/2^l)9 and Tetrahedral code π/2l|\pi/2^l\rangle0; its logical failure rate scales as π/2l|\pi/2^l\rangle1, with break-even against a physical π/2l|\pi/2^l\rangle2 gate at approximately π/2l|\pi/2^l\rangle3, while the π/2l|\pi/2^l\rangle4 variants achieve π/2l|\pi/2^l\rangle5 (Heußen et al., 2024). This suggests an MLTI interpretation in which transversal transfer between code levels substitutes for a dedicated magic-state factory.

A second related architecture is transversal STAR for neutral-atom simulation. There the inferred “multi-level” structure consists of four stages: transversal multi-rotation injection inside a patch, repeat-until-success teleportation to data patches, composition under correlated decoding with π/2l|\pi/2^l\rangle6 syndrome extraction between transversal Cliffords, and extension to high-rate codes with fold/swap-transversal Clifford layers (Ismail et al., 22 Sep 2025). The small-angle scaling reported for injection is

π/2l|\pi/2^l\rangle7

and the simulation-volume bound is

π/2l|\pi/2^l\rangle8

At π/2l|\pi/2^l\rangle9, the paper reports ll0 with approximately ll1 physical qubits, corresponding to a fully fault-tolerant computation requiring over ll2–ll3 ll4 gates (Ismail et al., 22 Sep 2025).

The qutrit literature supplies an additional variant. “Transversal AND in Quantum Codes” constructs a qutrit ll5 CSS code with a built-in transversal implementation of AND, derived from a symmetric T-depth-one compute–phase–uncompute circuit, and then concatenates it to a ll6 code preserving the same logical gate (Li et al., 4 Mar 2026). The paper does not name this MLTI, but its layered realization of non-Clifford action through code construction and concatenation is formally close to the quantum uses above.

4. Transversal injections in extremal combinatorics

In extremal combinatorics, MLTI denotes injective assignment across layers rather than state preparation. For a ll7-graph system

ll8

on a common ll9-vertex set kαβk|\alpha\rangle \to |\beta\rangle0, a kαβk|\alpha\rangle \to |\beta\rangle1-graph kαβk|\alpha\rangle \to |\beta\rangle2 is kαβk|\alpha\rangle \to |\beta\rangle3-transversal if there exists an injection

kαβk|\alpha\rangle \to |\beta\rangle4

such that kαβk|\alpha\rangle \to |\beta\rangle5 for all kαβk|\alpha\rangle \to |\beta\rangle6. In this setting, the injection itself is the transversal object (Cheng et al., 2021).

The main theorem for hypergraph systems states that for kαβk|\alpha\rangle \to |\beta\rangle7, kαβk|\alpha\rangle \to |\beta\rangle8, sufficiently large kαβk|\alpha\rangle \to |\beta\rangle9, and an +L|+_L\rangle0-vertex +L|+_L\rangle1-graph system +L|+_L\rangle2, if

+L|+_L\rangle3

for every +L|+_L\rangle4, then there exists an +L|+_L\rangle5-transversal tight Hamilton cycle (Cheng et al., 2021). The condition +L|+_L\rangle6 matches the fact that a tight Hamilton cycle has exactly +L|+_L\rangle7 edges, so an injection +L|+_L\rangle8 is exactly what is needed for one edge per layer. The proof uses the absorption method, with a transversal absorbing lemma, a connecting lemma, and a path-cover lemma, together with an auxiliary +L|+_L\rangle9-graph ZZ00, weak hypergraph regularity, reduced hypergraphs, and rainbow matching arguments that are then blown up to long transversal paths (Cheng et al., 2021).

A bipartite analogue appears in collections

ZZ01

on a common bipartition ZZ02 with ZZ03. If ZZ04 is a Hamiltonian path with ZZ05 and there exists an injection

ZZ06

such that each edge lies in its assigned level, then ZZ07 is a ZZ08-transversal isomorphic to a Hamiltonian path (Ma et al., 25 Jan 2026). Theorem 1.3 states that if

ZZ09

for each ZZ10, then either such a transversal Hamiltonian path exists or, when ZZ11 is even, all levels are the extremal disconnected graph

ZZ12

Theorem 1.4 raises the degree threshold to

ZZ13

and obtains Hamiltonian connectedness unless the odd-ZZ14 extremal family ZZ15 occurs (Ma et al., 25 Jan 2026). Here MLTI is equivalent to a full rainbow labeling of all ZZ16 edges of a spanning path.

These combinatorial uses are structurally strict: if the number of layers is below the number of required edges, full transversality is impossible. That exact impossibility statement is explicit for tight Hamilton cycles when ZZ17 and for bipartite Hamiltonian paths when fewer than ZZ18 layers are available (Cheng et al., 2021, Ma et al., 25 Jan 2026).

5. Layer-wise insertion in transformers and multi-level attacks on LLM systems

In transformer scaling, MLTI appears as the general strategy of inserting new transformer layers transversally across the stack at regular intervals, with Transformer Layer Injection (TLI) as the dense-transformer instantiation (Vo, 2024). If the base model has depth ZZ19, injection interval ZZ20, and injection set

ZZ21

then the new depth is

ZZ22

Each inserted block is initialized to be identity-preserving by duplicating the previous layer and zeroing the attention output projection and FFN down projection, so that the residual branch is initially zero. In pre-norm notation, if

ZZ23

TLI uses ZZ24 on the output projections, producing exact identity at initialization (Vo, 2024).

The training schedule is two-stage: first freeze the original ZZ25 layers and train only the injected ZZ26 layers, then optionally fine-tune all layers with LoRA or QLoRA. Reported experiments use LLama3 1B, 3B, and 8B on KoBEST and KMCQA. The paper reports lower initialization loss than DUS, fewer training steps to reach target performance, superior downstream accuracy, and several settings in which the injected models perform effectively even without additional training (Vo, 2024). Compared with Mixture of Experts, the architecture remains dense and avoids routing overhead; compared with DUS, the benefit is attributed to reduced initialization mismatch and preserved hidden-state continuity.

A different systems interpretation arises in LLM safety. “Transient Turn Injection” (TTI) is defined as a stateless multi-turn attack that distributes adversarial intent across isolated interactions, and the paper explicitly maps MLTI as a generalization across levels of an LLM-enabled pipeline: model front-end, tool calls, retrieval, external agents, content filters, and approval workflows (Rayhan et al., 23 Apr 2026). In the formalization given there, each level ZZ27 has a component ZZ28 and safeguard ZZ29, and the adversary seeks a final output ZZ30 while all local filters still return allow. Under this mapping, TTI is a subclass of MLTI specialized to turns at a single interface.

The evaluation reported for TTI shows substantial variation in robustness across models. In the cited run, Claude 3.5 Haiku has a safe response count of 49/50, GPT-4.1-mini and GPT-4o variants 46/50, while Gemini 2.0 Flash, Gemini 2.0 Flash-Lite, and Gemini 1.5 Flash range from 30/50 to 33/50 safe responses; TTI counts also exceed PAIR counts across models such as Gemini 2.0 Flash (PAIR=4, TTI=34) and GPT-4.1-mini (PAIR=2, TTI=8) (Rayhan et al., 23 Apr 2026). The mitigations proposed are session-level context aggregation, deep alignment, context-aware moderation, adaptive rate limiting, memory-based consistency checks, and continuous adversarial testing.

6. Shared structure, misconceptions, and open directions

A central misconception is that MLTI names a universally standardized object. The record is more heterogeneous. The term is formal and central in the logical rotation-state paper (Zhang et al., 28 Sep 2025); it is a named general strategy instantiated by TLI in transformer scaling (Vo, 2024); and in several quantum, combinatorial, and safety papers it is an editorial or inferred abstraction consistent with the original formalism rather than the paper’s own headline terminology (Gavriel et al., 2023, Heußen et al., 2024, Cheng et al., 2021, Rayhan et al., 23 Apr 2026). Any domain-independent definition must therefore be treated cautiously.

This suggests a shared structural template with three recurring features. First, there is a stratified object: code levels, hypergraph layers, transformer depth positions, or system components. Second, the injection is constrained: one edge per layer, one near-identity block per interval, one logical resource per level, or one locally admissible step per safeguard. Third, performance depends on preserving compatibility with the ambient structure: syndrome-consistent trajectories in FT protocols, codegree or minimum-degree thresholds in Hamiltonicity, hidden-state continuity in transformers, or stateless moderation gaps in LLM pipelines.

The open problems are likewise domain-specific. In explicit quantum MLTI, the paper points to further gains from integrating QLDPC codes and higher-rate architectures, and parameter search over ZZ31, ZZ32, ZZ33, and pumping fidelities remains architecture-dependent (Zhang et al., 28 Sep 2025). In surface-code and code-switching variants, the unresolved bottlenecks include exponential trajectory computation, efficient deterministic auxiliary-state preparation at higher distances, and adaptation to limited-connectivity hardware (Gavriel et al., 2022, Heußen et al., 2024). In hypergraph and bipartite transversality, open directions include loose cycles, alternative degree conditions, algorithmic tractability, and extensions to other numbers of levels (Cheng et al., 2021, Ma et al., 25 Jan 2026). In transformer scaling, practical questions concern faster specialized algorithms, ablations over ZZ34, and extension beyond depth-only scaling (Vo, 2024). In LLM safety, the stated priority is sequence-level and system-level policy reasoning rather than turn-local or component-local filtering, together with continuous automated red-teaming (Rayhan et al., 23 Apr 2026).

Across these literatures, MLTI functions less as a single theorem than as a reusable design principle: inject across levels, preserve a transversal constraint, and exploit the resulting structure to obtain fidelity, Hamiltonicity, scalability, or adversarial evasion, depending on the domain.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Multi-Level Transversal Injection (MLTI).