Multi-Level Transversal Injection (MLTI)
- MLTI is a protocol that prepares logical rotation states across multiple Clifford hierarchy levels using transversal injection and magic-state pumping.
- It alternates between error suppression and angle restoration to maintain fidelity while reducing physical error contributions.
- Its design principles extend to diverse fields, including quantum fault tolerance, combinatorial transversality, transformer scaling, and multi-turn LLM safety.
Multi-Level Transversal Injection (MLTI) most directly denotes a fault-tolerant protocol for preparing logical rotation states at arbitrary Clifford hierarchy levels by iterating logical-level transversal injection and magic-state pumping (Zhang et al., 28 Sep 2025). The same label, or closely related inferred abstractions, also appears in surface-code ancilla preparation, one-way-transversal code switching, transversal STAR gadgets, transformer depth upscaling, transversal Hamiltonicity, and stateless multi-turn attack models, where the common motif is a level-wise insertion or injective assignment constrained by a transversal structure (Gavriel et al., 2023, Heußen et al., 2024, Ismail et al., 22 Sep 2025, Vo, 2024, Cheng et al., 2021, Rayhan et al., 23 Apr 2026). This suggests that MLTI is not a single domain-independent construction, but a family of techniques built around structured injection across multiple layers, levels, or subsystems.
1. Terminology and scope
The expression “MLTI” is explicit in the 2025 letter on logical non-Clifford state preparation, where it names a concrete lattice-surgery-based method for preparing rotation states with overhead that decreases with Clifford hierarchy level and then plateaus (Zhang et al., 28 Sep 2025). In several other works, by contrast, the term is not part of the paper’s formal title or original notation; instead, it is introduced as an inferred abstraction consistent with the paper’s formalism, such as hierarchical transversal state preparation in stabilizer codes, multi-stage transversal STAR gadgets, or one-way-transversal code-switching pipelines (Gavriel et al., 2023, Gavriel et al., 2022, Ismail et al., 22 Sep 2025, Heußen et al., 2024).
| Context | Meaning of injection | Representative paper |
|---|---|---|
| Fault-tolerant quantum computing | Logical-level preparation or transfer of non-Clifford resources across code levels or patches | (Zhang et al., 28 Sep 2025) |
| Extremal combinatorics | Injection assigning edges of a spanning structure to distinct layers | (Cheng et al., 2021) |
| Transformer scaling | Regular insertion of new layers across depth with near-identity initialization | (Vo, 2024) |
| LLM safety | Distribution of adversarial intent across turns or system levels | (Rayhan et al., 23 Apr 2026) |
Within quantum computing, “transversal” retains its standard FT connotation: operations are arranged so that error propagation is constrained by code structure, while the injected object is typically a resource state, a non-Clifford phase, or a temporary encoding transfer. Within combinatorics, “transversal” refers to selecting one edge from each layer, or more precisely assigning distinct edges to distinct layers via an injection. In transformer scaling, the term denotes layer insertion across the stack rather than end-appending. In LLM safety, it denotes adversarial traversal across levels of a pipeline while each local safeguard still returns allow.
2. Explicit MLTI for logical rotation-state preparation
In its explicit quantum-information sense, MLTI is a protocol for preparing rotation states at arbitrary Clifford hierarchy levels. The basic single-qubit objects are the -axis rotation
and the corresponding rotation state
A level- rotation is , and the associated state is a level- rotation state (Zhang et al., 28 Sep 2025).
The core physical-level primitive is transversal injection of the form . Starting from , one applies 0 to each of the 1 physical qubits in 2. In the noiseless case, post-selection yields
3
with
4
Under circuit-level noise with physical error rate 5, the output infidelity scales as
6
MLTI lifts this primitive to the logical level by using lattice surgery to measure 7 across 8 logical patches, thereby implementing the same ideal map 9 on logical inputs (Zhang et al., 28 Sep 2025).
The central suppression statement is the per-level theorem:
0
where 1 is the input infidelity to 2. This is the mechanism by which successive logical levels reduce error. Because repeated injection would otherwise shrink the angle too aggressively, MLTI introduces magic-state pumping: after each level, apply 3 and 4 so that the next level’s input angle is brought close to 5 in magnitude. The protocol therefore alternates suppression and angle restoration (Zhang et al., 28 Sep 2025).
The implementation is surface-code based and uses asymmetric rotated patches with 6, since a 7 error sends 8 to 9 and therefore dominates infidelity, whereas 0 errors contribute only 1. The logical error estimates used in the resource model are
2
with
3
and for rectangular patches the paper rescales by area ratios:
4
This asymmetry is part of the overhead reduction, not a secondary optimization (Zhang et al., 28 Sep 2025).
A second technical contribution is elimination of off-diagonal terms “for free.” Naïve dephasing in the 5 basis would require an extra 6, but MLTI integrates dephasing into teleportation by randomly pre-applying 7 to the ancilla with probability 8 and switching between two teleportation channels, 9 and 0. This cancels off-diagonal contributions without introducing extra rotations, thereby making infidelity and trace distance coincide for the prepared ancilla (Zhang et al., 28 Sep 2025).
The resource model is stated in space-time volume, defined as physical-qubit count times QEC cycles per successfully produced target state. For the first level,
1
For pumping,
2
For the second level,
3
Quantitatively, at 4 and target infidelity 5, the reported space-time volume is below state distillation for all 6, and for 7 it is lower by a factor 8; with up to MLTI level 9, high-fidelity preparation of all rotation states with 0 is achievable (Zhang et al., 28 Sep 2025).
3. Quantum antecedents and related architectures
MLTI in the 2025 sense sits within a broader line of transversal-injection ideas. In surface-code state preparation, Transversal Injection (TI) initializes every data qubit in
1
before standard stabilizer measurements. The state 2 is projected by the stabilizer trajectory into a logical non-Pauli state, and the logical amplitudes are computed by trajectory-dependent sums of monomials 3 (Gavriel et al., 2023). A related stabilizer-code treatment states the same mechanism in terms of uniform physical rotations, trajectory-conditioned amplitude sums, and gate teleportation of 4 or 5 resource states; its hierarchical extension is described there as a multi-level application of TI across concatenated levels (Gavriel et al., 2022). In both accounts, the preparation is probabilistic and heralded, and the logical output depends on measured syndrome history rather than being fixed a priori.
Other FT constructions instantiate the same multi-stage pattern without using the term explicitly. One-way-transversal code switching realizes a logical 6 gate by teleporting from a self-dual code to a triorthogonal code, applying a transversal 7, and teleporting back using only transversal CNOTs, transversal single-qubit rotations, and transversal measurements (Heußen et al., 2024). For the 8 pair, the protocol uses the Steane code 9 and Tetrahedral code 0; its logical failure rate scales as 1, with break-even against a physical 2 gate at approximately 3, while the 4 variants achieve 5 (Heußen et al., 2024). This suggests an MLTI interpretation in which transversal transfer between code levels substitutes for a dedicated magic-state factory.
A second related architecture is transversal STAR for neutral-atom simulation. There the inferred “multi-level” structure consists of four stages: transversal multi-rotation injection inside a patch, repeat-until-success teleportation to data patches, composition under correlated decoding with 6 syndrome extraction between transversal Cliffords, and extension to high-rate codes with fold/swap-transversal Clifford layers (Ismail et al., 22 Sep 2025). The small-angle scaling reported for injection is
7
and the simulation-volume bound is
8
At 9, the paper reports 0 with approximately 1 physical qubits, corresponding to a fully fault-tolerant computation requiring over 2–3 4 gates (Ismail et al., 22 Sep 2025).
The qutrit literature supplies an additional variant. “Transversal AND in Quantum Codes” constructs a qutrit 5 CSS code with a built-in transversal implementation of AND, derived from a symmetric T-depth-one compute–phase–uncompute circuit, and then concatenates it to a 6 code preserving the same logical gate (Li et al., 4 Mar 2026). The paper does not name this MLTI, but its layered realization of non-Clifford action through code construction and concatenation is formally close to the quantum uses above.
4. Transversal injections in extremal combinatorics
In extremal combinatorics, MLTI denotes injective assignment across layers rather than state preparation. For a 7-graph system
8
on a common 9-vertex set 0, a 1-graph 2 is 3-transversal if there exists an injection
4
such that 5 for all 6. In this setting, the injection itself is the transversal object (Cheng et al., 2021).
The main theorem for hypergraph systems states that for 7, 8, sufficiently large 9, and an 0-vertex 1-graph system 2, if
3
for every 4, then there exists an 5-transversal tight Hamilton cycle (Cheng et al., 2021). The condition 6 matches the fact that a tight Hamilton cycle has exactly 7 edges, so an injection 8 is exactly what is needed for one edge per layer. The proof uses the absorption method, with a transversal absorbing lemma, a connecting lemma, and a path-cover lemma, together with an auxiliary 9-graph 00, weak hypergraph regularity, reduced hypergraphs, and rainbow matching arguments that are then blown up to long transversal paths (Cheng et al., 2021).
A bipartite analogue appears in collections
01
on a common bipartition 02 with 03. If 04 is a Hamiltonian path with 05 and there exists an injection
06
such that each edge lies in its assigned level, then 07 is a 08-transversal isomorphic to a Hamiltonian path (Ma et al., 25 Jan 2026). Theorem 1.3 states that if
09
for each 10, then either such a transversal Hamiltonian path exists or, when 11 is even, all levels are the extremal disconnected graph
12
Theorem 1.4 raises the degree threshold to
13
and obtains Hamiltonian connectedness unless the odd-14 extremal family 15 occurs (Ma et al., 25 Jan 2026). Here MLTI is equivalent to a full rainbow labeling of all 16 edges of a spanning path.
These combinatorial uses are structurally strict: if the number of layers is below the number of required edges, full transversality is impossible. That exact impossibility statement is explicit for tight Hamilton cycles when 17 and for bipartite Hamiltonian paths when fewer than 18 layers are available (Cheng et al., 2021, Ma et al., 25 Jan 2026).
5. Layer-wise insertion in transformers and multi-level attacks on LLM systems
In transformer scaling, MLTI appears as the general strategy of inserting new transformer layers transversally across the stack at regular intervals, with Transformer Layer Injection (TLI) as the dense-transformer instantiation (Vo, 2024). If the base model has depth 19, injection interval 20, and injection set
21
then the new depth is
22
Each inserted block is initialized to be identity-preserving by duplicating the previous layer and zeroing the attention output projection and FFN down projection, so that the residual branch is initially zero. In pre-norm notation, if
23
TLI uses 24 on the output projections, producing exact identity at initialization (Vo, 2024).
The training schedule is two-stage: first freeze the original 25 layers and train only the injected 26 layers, then optionally fine-tune all layers with LoRA or QLoRA. Reported experiments use LLama3 1B, 3B, and 8B on KoBEST and KMCQA. The paper reports lower initialization loss than DUS, fewer training steps to reach target performance, superior downstream accuracy, and several settings in which the injected models perform effectively even without additional training (Vo, 2024). Compared with Mixture of Experts, the architecture remains dense and avoids routing overhead; compared with DUS, the benefit is attributed to reduced initialization mismatch and preserved hidden-state continuity.
A different systems interpretation arises in LLM safety. “Transient Turn Injection” (TTI) is defined as a stateless multi-turn attack that distributes adversarial intent across isolated interactions, and the paper explicitly maps MLTI as a generalization across levels of an LLM-enabled pipeline: model front-end, tool calls, retrieval, external agents, content filters, and approval workflows (Rayhan et al., 23 Apr 2026). In the formalization given there, each level 27 has a component 28 and safeguard 29, and the adversary seeks a final output 30 while all local filters still return allow. Under this mapping, TTI is a subclass of MLTI specialized to turns at a single interface.
The evaluation reported for TTI shows substantial variation in robustness across models. In the cited run, Claude 3.5 Haiku has a safe response count of 49/50, GPT-4.1-mini and GPT-4o variants 46/50, while Gemini 2.0 Flash, Gemini 2.0 Flash-Lite, and Gemini 1.5 Flash range from 30/50 to 33/50 safe responses; TTI counts also exceed PAIR counts across models such as Gemini 2.0 Flash (PAIR=4, TTI=34) and GPT-4.1-mini (PAIR=2, TTI=8) (Rayhan et al., 23 Apr 2026). The mitigations proposed are session-level context aggregation, deep alignment, context-aware moderation, adaptive rate limiting, memory-based consistency checks, and continuous adversarial testing.
6. Shared structure, misconceptions, and open directions
A central misconception is that MLTI names a universally standardized object. The record is more heterogeneous. The term is formal and central in the logical rotation-state paper (Zhang et al., 28 Sep 2025); it is a named general strategy instantiated by TLI in transformer scaling (Vo, 2024); and in several quantum, combinatorial, and safety papers it is an editorial or inferred abstraction consistent with the original formalism rather than the paper’s own headline terminology (Gavriel et al., 2023, Heußen et al., 2024, Cheng et al., 2021, Rayhan et al., 23 Apr 2026). Any domain-independent definition must therefore be treated cautiously.
This suggests a shared structural template with three recurring features. First, there is a stratified object: code levels, hypergraph layers, transformer depth positions, or system components. Second, the injection is constrained: one edge per layer, one near-identity block per interval, one logical resource per level, or one locally admissible step per safeguard. Third, performance depends on preserving compatibility with the ambient structure: syndrome-consistent trajectories in FT protocols, codegree or minimum-degree thresholds in Hamiltonicity, hidden-state continuity in transformers, or stateless moderation gaps in LLM pipelines.
The open problems are likewise domain-specific. In explicit quantum MLTI, the paper points to further gains from integrating QLDPC codes and higher-rate architectures, and parameter search over 31, 32, 33, and pumping fidelities remains architecture-dependent (Zhang et al., 28 Sep 2025). In surface-code and code-switching variants, the unresolved bottlenecks include exponential trajectory computation, efficient deterministic auxiliary-state preparation at higher distances, and adaptation to limited-connectivity hardware (Gavriel et al., 2022, Heußen et al., 2024). In hypergraph and bipartite transversality, open directions include loose cycles, alternative degree conditions, algorithmic tractability, and extensions to other numbers of levels (Cheng et al., 2021, Ma et al., 25 Jan 2026). In transformer scaling, practical questions concern faster specialized algorithms, ablations over 34, and extension beyond depth-only scaling (Vo, 2024). In LLM safety, the stated priority is sequence-level and system-level policy reasoning rather than turn-local or component-local filtering, together with continuous automated red-teaming (Rayhan et al., 23 Apr 2026).
Across these literatures, MLTI functions less as a single theorem than as a reusable design principle: inject across levels, preserve a transversal constraint, and exploit the resulting structure to obtain fidelity, Hamiltonicity, scalability, or adversarial evasion, depending on the domain.