Papers
Topics
Authors
Recent
Search
2000 character limit reached

Diagonal Block-Encodings Overview

Updated 14 July 2026
  • Diagonal block-encodings are structured block-encoding methods that represent operators diagonal in the computational basis, facilitating efficient quantum algorithm design.
  • They employ techniques such as the linear-combination-of-unitaries and oracle-based constructions to decompose operators, enabling applications in QSVT, PDE simulation, and linear system solvers.
  • Recent developments extend these methods to finite-difference Poisson, periodic, and semiseparable matrices, optimizing circuit depth, gate counts, and error management in practical implementations.

Searching arXiv for papers on diagonal block-encodings and related structured block-encoding methods. Diagonal block-encodings are block-encodings in which the target operator is diagonal in the computational basis, D=j=0N1fjjjD=\sum_{j=0}^{N-1} f_j\,|j\rangle\langle j|, or, in closely related structured settings, a matrix is first reduced to diagonal or Pauli-block-diagonal components and then embedded into a larger unitary. In the standard formulation, an (α,a,ε)(\alpha,a,\varepsilon)-block-encoding of an N×NN\times N operator AA is an (n+a)(n+a)-qubit unitary UU satisfying

(0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,

with N=2nN=2^n. Recent work treats diagonal block-encodings as a basic primitive for QSVT, linear-system algorithms, PDE simulation, and parameterized scientific computing, while related constructions extend the same design logic to periodic diagonal matrices, finite-difference Poisson operators, and semiseparable factorizations (Yano et al., 2 Mar 2026, Ty et al., 2024).

1. Formal definition and LCU realization

For a diagonal operator DD, the defining relation specializes to

(0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.

Equivalently, for any (α,a,ε)(\alpha,a,\varepsilon)0,

(α,a,ε)(\alpha,a,\varepsilon)1

so the block-encoding exposes diagonal data directly in basis-state amplitudes (Yano et al., 2 Mar 2026).

A standard exact realization uses the linear-combination-of-unitaries framework. If

(α,a,ε)(\alpha,a,\varepsilon)2

one defines a state-preparation unitary (α,a,ε)(\alpha,a,\varepsilon)3 and the controlled unitary (α,a,ε)(\alpha,a,\varepsilon)4, yielding

(α,a,ε)(\alpha,a,\varepsilon)5

In the finite-difference Poisson constructions, (α,a,ε)(\alpha,a,\varepsilon)6 for all (α,a,ε)(\alpha,a,\varepsilon)7, so the subnormalized factor is (α,a,ε)(\alpha,a,\varepsilon)8, and unitarity is inherited directly from the controlled-unitary building blocks and the Hadamard-tree preparation/unpreparation on the control register (Ty et al., 2024).

This formalism makes diagonal block-encoding a special case of structured LCU: the operator can be diagonal from the outset, or it can be decomposed into a small number of diagonal, shift, or Pauli-block terms. The latter viewpoint underlies most of the recent extensions.

2. Direct constructions for diagonal operators

One line of work constructs diagonal block-encodings from quantum-accessible classical data. Three loading models are used. A probability oracle (α,a,ε)(\alpha,a,\varepsilon)9 yields a N×NN\times N0 diagonal block-encoding of N×NN\times N1 with N×NN\times N2. A phase oracle N×NN\times N3 is itself a N×NN\times N4-block-encoding of N×NN\times N5. A binary oracle N×NN\times N6 loads a fixed-point expansion N×NN\times N7 and is then used to control subsequent gates rather than serving as a block-encoding on its own (Yano et al., 2 Mar 2026).

For spatially varying coefficients in PDEs, the same paper considers analytic approximations of the form

N×NN\times N8

which induces the LCU representation

N×NN\times N9

The implementation uses the standard PREP/SEL/UNPREP pattern: PREP prepares AA0, SEL applies AA1, and UNPREP inverts PREP. The resulting unitary uses AA2 ancillas (Yano et al., 2 Mar 2026).

A distinct but closely related construction addresses periodic diagonal structure through an explicit phase gadget. For

AA3

binary factorization gives

AA4

Using one ancilla, Hadamards, controlled-AA5, and AA6, the cosine operator satisfies

AA7

so the AA8 ancilla subspace carries the desired cosine block. In the special case AA9 with (n+a)(n+a)0, the circuit uses one ancilla and has

(n+a)(n+a)1

hence (n+a)(n+a)2 (Zecchi et al., 11 Feb 2026).

These direct constructions illustrate two complementary paradigms: oracle-based diagonal loading and analytically structured phase synthesis. Both avoid unstructured dense-matrix input models, but they differ in what must be supplied classically: either a query model for (n+a)(n+a)3 or an explicit harmonic/phase representation.

3. Extension from diagonal operators to structured matrices

The most prominent extension is the finite-difference Poisson setting. The matrix family

(n+a)(n+a)4

arises from centered three-point finite-difference discretizations of Poisson’s equation on the (n+a)(n+a)5-dimensional unit hypercube (n+a)(n+a)6 with periodic, Dirichlet, Neumann, or Robin boundary conditions. Because each one-dimensional stencil matrix is circulant or nearly circulant, it admits a combinatorial block-diagonalization of the form

(n+a)(n+a)7

Here (n+a)(n+a)8 for periodic, Dirichlet, and Neumann cases, and (n+a)(n+a)9 for Robin boundary conditions. In the periodic case,

UU0

with UU1 and UU2. The periodic, Dirichlet, and Neumann cases require UU3 LCU terms, while the Robin case needs UU4 nonzero terms, padded to UU5 by two zero operators in the LCU (Ty et al., 2024).

The corresponding circuits retain a diagonal-block flavor because each controlled unitary is elementary: UU6, UU7 on the least-significant qubit, UU8, or, in the Robin case, controlled-UU9 rotations on the most-significant qubit plus (0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,0 and its inverse. The incrementer

(0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,1

is implemented from a carry-lookahead adder (Ty et al., 2024).

A different structured extension appears for one-pair semiseparable matrices. A real symmetric one-pair semiseparable matrix (0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,2 generated by vectors (0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,3 and (0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,4 satisfies

(0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,5

and admits the factorization

(0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,6

with (0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,7 diagonal, (0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,8 the lower-triangular all-ones matrix, and (0aIN)U(0aIN)=Aα±ε,\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,9 diagonal. The key subroutine is an N=2nN=2^n0-block-encoding of an arbitrary diagonal matrix N=2nN=2^n1, built from an oracle N=2nN=2^n2, a reversible Boolean circuit computing a fixed-point N=2nN=2^n3, and a N=2nN=2^n4 gadget of controlled N=2nN=2^n5-rotations plus a controlled-N=2nN=2^n6 for the sign. The full semiseparable construction uses N=2nN=2^n7 ancillary qubits and polylogarithmic depth (Antonioli et al., 19 Mar 2026).

Taken together, these results suggest a broader structural interpretation of diagonal block-encoding: the diagonal primitive is not limited to already diagonal matrices, but also serves as the central component in factorizations or decompositions of matrices with periodic, semiseparable, or nearly circulant structure.

4. Circuit depth, gate counts, and synthesis optimization

Reported resource statements vary sharply with the structure being exploited.

Construction Reported resources
Simple diagonal N=2nN=2^n8 N=2nN=2^n9
Banded periodic DD0 DD1 with naive shift, DD2 with optimized DD3 adder
Finite-difference Poisson matrices ancillas DD4, Toffoli-depth DD5, overall gate count DD6
One-pair semiseparable matrices DD7 ancillas, polylogarithmic depth, error DD8

For the Poisson matrices, the key asymptotic statement is

DD9

because the only nontrivial subroutines are multi-controlled Pauli gates of (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.0 depth and the incrementer (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.1 of depth (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.2, with (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.3. The paper further states ancilla qubits (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.4 for the LCU plus (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.5 for arithmetic, overall gate count (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.6, and compares this against an “oracle” approach of polylog (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.7, poly (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.8 depth and a classical sparse-matrix encode of depth (0aIN)U(0aIN)=Dα.(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.9 (Ty et al., 2024).

For periodic diagonal structure, the simple diagonal case is linear in (α,a,ε)(\alpha,a,\varepsilon)00, whereas the full banded periodic matrix depends on how shifts are implemented. PREP and PREP(α,a,ε)(\alpha,a,\varepsilon)01 on two ancillas cost (α,a,ε)(\alpha,a,\varepsilon)02 two-qubit gates if implemented by small fixed unitaries. The SELECT stage is dominated either by naive increment/decrement circuits of (α,a,ε)(\alpha,a,\varepsilon)03 Toffoli cost or by optimized (α,a,ε)(\alpha,a,\varepsilon)04 adders, leading to

(α,a,ε)(\alpha,a,\varepsilon)05

This separates the genuinely diagonal part from the shift overhead (Zecchi et al., 11 Feb 2026).

A separate synthesis result improves generic block-encoding circuits through “diagonal matrix migration.” In the single-ancilla protocol for (α,a,ε)(\alpha,a,\varepsilon)06 matrices, a data-qubit diagonal operator commutes through a uniformly controlled (α,a,ε)(\alpha,a,\varepsilon)07 layer, allowing diagonal corrections to be absorbed into subsequent stages of a block-ZXZ recursion. The resulting leading C-NOT count for general block-encoding is

(α,a,ε)(\alpha,a,\varepsilon)08

and for a rank-(α,a,ε)(\alpha,a,\varepsilon)09 matrix,

(α,a,ε)(\alpha,a,\varepsilon)10

The paper also states the counting lower bound

(α,a,ε)(\alpha,a,\varepsilon)11

for a general block-encoding of a (α,a,ε)(\alpha,a,\varepsilon)12 complex matrix (Li et al., 17 Mar 2026).

The resource picture is therefore heterogeneous: diagonality can collapse costs to (α,a,ε)(\alpha,a,\varepsilon)13, but neighboring structural requirements such as shifts, factor compositions, or generic state preparation may dominate the actual circuit.

5. Role in PDE simulation, QSVT, and parameterized scientific computing

Diagonal block-encodings are especially prominent in PDE-oriented quantum algorithms because discretized coefficient fields often appear as diagonal matrices. In the parameterized PDE framework, diagonal coefficient operators (α,a,ε)(\alpha,a,\varepsilon)14 are combined with finite-difference block-encodings (α,a,ε)(\alpha,a,\varepsilon)15 to obtain block-encodings of first-order and second-order PDE generators. Under a Fourier assumption for the coefficients, the resulting gate count is

(α,a,ε)(\alpha,a,\varepsilon)16

with ancilla cost (α,a,ε)(\alpha,a,\varepsilon)17 for the diagonal LCU-Fourier part and error (α,a,ε)(\alpha,a,\varepsilon)18 controlled by Fourier truncation (Yano et al., 2 Mar 2026).

Once a block-encoding (α,a,ε)(\alpha,a,\varepsilon)19 of (α,a,ε)(\alpha,a,\varepsilon)20 is available, QSVT is used to implement (α,a,ε)(\alpha,a,\varepsilon)21 or another polynomial approximation on the singular values. The corresponding sequence length is

(α,a,ε)(\alpha,a,\varepsilon)22

The same framework extends to parameter-dependent operators

(α,a,ε)(\alpha,a,\varepsilon)23

so that a design register can coherently control PDE evolution and objective evaluation in PDE-constrained optimization (Yano et al., 2 Mar 2026).

The detailed two-dimensional wave-equation example uses a shifted Gaussian speed profile

(α,a,ε)(\alpha,a,\varepsilon)24

approximated by a truncated Fourier series of order (α,a,ε)(\alpha,a,\varepsilon)25. The induced coefficient operator (α,a,ε)(\alpha,a,\varepsilon)26 is block-encoded with (α,a,ε)(\alpha,a,\varepsilon)27 parameters. After adding another 6 ancillas for the (α,a,ε)(\alpha,a,\varepsilon)28 system selector and finite-difference block-encodings, the forward simulation uses 20 qubits; turning on the (α,a,ε)(\alpha,a,\varepsilon)29 registers yields 28 total qubits. The reported numerical simulations reproduce wave-front slowing near the Gaussian bump and show an objective optimum near (α,a,ε)(\alpha,a,\varepsilon)30 (Yano et al., 2 Mar 2026).

The same diagonal and near-diagonal methodology appears in other PDE-related settings. For the periodic-diagonal banded matrix

(α,a,ε)(\alpha,a,\varepsilon)31

the block-encoding is intended for QSVT tasks such as matrix inversion and matrix exponentiation, with advection-diffusion-reaction dynamics given as an application (Zecchi et al., 11 Feb 2026). For finite-difference Poisson systems, combining the (α,a,ε)(\alpha,a,\varepsilon)32-depth block-encoding with the optimal adiabatic quantum linear solver of Costa et al. yields overall wall-clock depth

(α,a,ε)(\alpha,a,\varepsilon)33

for preparing a quantum state representation of the PDE solution (Ty et al., 2024).

6. Error models, normalization, and implementation trade-offs

A central distinction in the literature is between exact and approximate diagonal block-encodings. The finite-difference Poisson constructions are exact at the block-encoding level because all (α,a,ε)(\alpha,a,\varepsilon)34 are exactly unitary and the LCU coefficients are uniform. By contrast, Fourier-based PDE coefficient encodings incur truncation error (α,a,ε)(\alpha,a,\varepsilon)35, and semiseparable encodings accumulate approximation error from the oracle-loaded diagonal data, the fixed-point (α,a,ε)(\alpha,a,\varepsilon)36 circuit, and the product of multiple block-encodings (Ty et al., 2024, Yano et al., 2 Mar 2026, Antonioli et al., 19 Mar 2026).

Normalization can also be decisive. In the semiseparable case, the scaling factor is (α,a,ε)(\alpha,a,\varepsilon)37, and the final error bound is

(α,a,ε)(\alpha,a,\varepsilon)38

This means that polylogarithmic depth does not by itself imply small effective query complexity, because normalization and approximation constants still enter higher-level algorithms (Antonioli et al., 19 Mar 2026).

A separate trade-off appears in fault-tolerant implementations of diagonal quadratic operators such as

(α,a,ε)(\alpha,a,\varepsilon)39

In the LCU/block-encoding setting, the qubit signed-binary baseline uses projectors (α,a,ε)(\alpha,a,\varepsilon)40 with per-call cost

(α,a,ε)(\alpha,a,\varepsilon)41

The native qudit LCU expands (α,a,ε)(\alpha,a,\varepsilon)42 in generalized Pauli powers (α,a,ε)(\alpha,a,\varepsilon)43, with PREP and SELECT requiring (α,a,ε)(\alpha,a,\varepsilon)44 embedded two-level rotations in the fixed-encoding model. The paper states that the qubit encoding is asymptotically cheaper in (α,a,ε)(\alpha,a,\varepsilon)45, while finite-(α,a,ε)(\alpha,a,\varepsilon)46 threshold analysis identifies regions where qudits can offer constant-factor savings: at (α,a,ε)(\alpha,a,\varepsilon)47, (α,a,ε)(\alpha,a,\varepsilon)48 only for (α,a,ε)(\alpha,a,\varepsilon)49 in the precision-dominated regime (α,a,ε)(\alpha,a,\varepsilon)50, and up to prime (α,a,ε)(\alpha,a,\varepsilon)51 in the time-dominated regime (α,a,ε)(\alpha,a,\varepsilon)52 (Godwood et al., 29 Apr 2026).

These results indicate that diagonal structure alone does not fix the full cost profile. The dominant overhead can lie in PREP, in shifts or adders, in normalization, in approximation error, or in fault-tolerant rotation synthesis. The common advantage of diagonal block-encodings is therefore not the elimination of all complexity, but the conversion of otherwise unstructured matrix input into a form compatible with explicit circuit analysis, QSVT composition, and structure-dependent asymptotic improvements.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Diagonal Block-Encodings.