---
title: Diagonal Block-Encodings Overview
url: https://www.emergentmind.com/topics/diagonal-block-encodings
type: topic
---

# Diagonal Block-Encodings Overview

Searching arXiv for recent papers on diagonal block-encodings and related structured block-encoding methods.
Diagonal block-encodings are block-encodings in which the target operator is diagonal in the computational basis, \(D=\sum_{j=0}^{N-1} f_j\,|j\rangle\langle j|\), or, in closely related structured settings, a matrix is first reduced to diagonal or Pauli-block-diagonal components and then embedded into a larger unitary. In the standard formulation, an \((\alpha,a,\varepsilon)\)-block-encoding of an \(N\times N\) operator \(A\) is an \((n+a)\)-qubit unitary \(U\) satisfying
\[
\bigl(\langle 0|^{\otimes a}\otimes I_N\bigr)\,U\,\bigl(|0\rangle^{\otimes a}\otimes I_N\bigr)=\frac{A}{\alpha}\pm \varepsilon,
\]
with \(N=2^n\). Recent work treats diagonal block-encodings as a basic primitive for QSVT, linear-system algorithms, PDE simulation, and parameterized scientific computing, while related constructions extend the same design logic to periodic diagonal matrices, finite-difference Poisson operators, and semiseparable factorizations [2603.01358], [2410.05241].

## 1. Formal definition and LCU realization

For a diagonal operator \(D\), the defining relation specializes to
\[
(\langle0|^{\otimes a}\otimes I_N)\,U\,(|0\rangle^{\otimes a}\otimes I_N)=\frac{D}{\alpha}.
\]
Equivalently, for any \(|\psi\rangle\in\mathbb C^N\),
\[
(\langle0|^{\otimes a}\langle\psi|)\,U\,(|0\rangle^{\otimes a}|\psi\rangle)
=\frac{1}{\alpha}\sum_j f_j\,|\langle j|\psi\rangle|^2,
\]
so the block-encoding exposes diagonal data directly in basis-state amplitudes [2603.01358].

A standard exact realization uses the linear-combination-of-unitaries framework. If
\[
\mathbf L=\sum_{j=1}^m \alpha_j U_j,\qquad U_j^\dagger U_j=I,
\]
one defines a state-preparation unitary \(W\) and the controlled unitary \(U_{\rm ctrl}=\sum_j |j\rangle\langle j|\otimes U_j\), yielding
\[
\bigl(\bra{0}_{\rm ctrl}\otimes I\bigr)
\bigl(W^\dagger\otimes I\bigr)
U_{\rm ctrl}
\bigl(W\otimes I\bigr)
\bigl(\ket{0}_{\rm ctrl}\otimes I\bigr)
=
\frac{\mathbf L}{\sum_j \alpha_j}.
\]
In the finite-difference Poisson constructions, \(\alpha_j=1\) for all \(j\), so the subnormalized factor is \(\eta=m\), and unitarity is inherited directly from the controlled-unitary building blocks and the Hadamard-tree preparation/unpreparation on the control register [2410.05241].

This formalism makes diagonal block-encoding a special case of structured LCU: the operator can be diagonal from the outset, or it can be decomposed into a small number of diagonal, shift, or Pauli-block terms. The latter viewpoint underlies most of the recent extensions.

## 2. Direct constructions for diagonal operators

One line of work constructs diagonal block-encodings from quantum-accessible classical data. Three loading models are used. A probability oracle \(U_{\rm prob}\) yields a \((1,a,0)\) diagonal block-encoding of \(\sum_\xi g(\xi)|\xi\rangle\langle\xi|\) with \(g(\xi)=2p(\xi)-1\). A phase oracle \(U_{\rm phase}\) is itself a \((1,a,0)\)-block-encoding of \(\sum_\xi e^{i g(\xi)}|\xi\rangle\langle\xi|\). A binary oracle \(U_{\rm binary}\) loads a fixed-point expansion \(\tilde g(\xi)\) and is then used to control subsequent gates rather than serving as a block-encoding on its own [2603.01358].

For spatially varying coefficients in PDEs, the same paper considers analytic approximations of the form
\[
f(\mathbf x^{[j]})\approx \sum_{k\in\mathcal K} c_k\,e^{i\pi k\cdot \mathbf x^{[j]}},
\]
which induces the LCU representation
\[
D\approx \sum_{k\in\mathcal K} c_k\,U_x^k,\qquad
U_x^k:|j\rangle\mapsto e^{i\pi k j/(N-1)}\,|j\rangle.
\]
The implementation uses the standard PREP/SEL/UNPREP pattern: PREP prepares \(\sum_k \sqrt{|c_k|/\alpha}\,|k\rangle\), SEL applies \(\sum_k |k\rangle\langle k|\otimes U_x^k\), and UNPREP inverts PREP. The resulting unitary uses \(\lceil\log|\mathcal K|\rceil\) ancillas [2603.01358].

A distinct but closely related construction addresses periodic diagonal structure through an explicit phase gadget. For
\[
V(\omega)=\mathrm{diag}\bigl(1,e^{i\omega},e^{i2\omega},\dots,e^{i(N-1)\omega}\bigr),
\]
binary factorization gives
\[
V(\omega)=\prod_{q=0}^{n-1} P(2^q\omega),\qquad
P(\theta)=\begin{pmatrix}1&0\\0&e^{i\theta}\end{pmatrix}.
\]
Using one ancilla, Hadamards, controlled-\(V(2\omega)\), and \(V(\omega)^\dagger\), the cosine operator satisfies
\[
U_{C(\omega)}\bigl(|0\rangle|\psi\rangle\bigr)
=
|0\rangle\,C(\omega)|\psi\rangle
+i\,|1\rangle\,S(\omega)|\psi\rangle,
\]
so the \(|0\rangle\) ancilla subspace carries the desired cosine block. In the special case \(D=\mathrm{diag}(d_0,\dots,d_{N-1})\) with \(d_k=\cos(k\omega)\), the circuit uses one ancilla and has
\[
G_{\rm diag}(n)=3n\;P\text{-gates}+2n\;\text{CNOTs}+2\;H\text{'s},
\]
hence \(G_{\rm diag}(n)=\Theta(n)\) [2602.10589].

These direct constructions illustrate two complementary paradigms: oracle-based diagonal loading and analytically structured phase synthesis. Both avoid unstructured dense-matrix input models, but they differ in what must be supplied classically: either a query model for \(f_j\) or an explicit harmonic/phase representation.

## 3. Extension from diagonal operators to structured matrices

The most prominent extension is the finite-difference Poisson setting. The matrix family
\[
\mathbf{L}\in\{\mathbf{L}_p,\mathbf{L}_D,\mathbf{L}_N,\mathbf{L}_R\}
\]
arises from centered three-point finite-difference discretizations of Poisson’s equation on the \(d\)-dimensional unit hypercube \([0,1]^d\) with periodic, Dirichlet, Neumann, or Robin boundary conditions. Because each one-dimensional stencil matrix is circulant or nearly circulant, it admits a combinatorial block-diagonalization of the form
\[
\mathbf L=\sum_{c\in\{0,1\}}\sum_{\hat\sigma\in\mathcal S}
\chi_{\hat\sigma}\,\mathbf L_{c,\hat\sigma},
\qquad
\mathbf L_{c,\hat\sigma}
=
\mathbf P_c\,
\bigl[\mathrm{diag}(\alpha_0,\dots,\alpha_{N/2-1})\otimes \hat\sigma\bigr]\,
\mathbf P_c^\dagger.
\]
Here \(\mathcal S=\{\mathbf I,\mathbf X\}\) for periodic, Dirichlet, and Neumann cases, and \(\mathcal S=\{\mathbf I,\mathbf X,\mathbf Z\}\) for Robin boundary conditions. In the periodic case,
\[
\mathbf{L}_p
=
\sum_{c=0}^1\sum_{\hat\sigma\in\{\mathbf I,\mathbf X\}}
\chi_{\hat\sigma}\,
\mathbf P_c\bigl[I_{N/2}\otimes\hat\sigma\bigr]\mathbf P_c^\dagger,
\]
with \(\chi_{\mathbf I}=+1\) and \(\chi_{\mathbf X}=-1\). The periodic, Dirichlet, and Neumann cases require \(m=4\) LCU terms, while the Robin case needs \(m=6\) nonzero terms, padded to \(8\) by two zero operators in the LCU [2410.05241].

The corresponding circuits retain a diagonal-block flavor because each controlled unitary is elementary: \(I^{\otimes n}\), \(X\) on the least-significant qubit, \(\mathrm{ADD1}(I^{\otimes n-1}\otimes X)\mathrm{ADD1}^\dagger\), or, in the Robin case, controlled-\(Z\) rotations on the most-significant qubit plus \(\mathrm{ADD1}\) and its inverse. The incrementer
\[
\mathrm{ADD1}\equiv \sum_k |k+1\!\!\!\!\pmod N\rangle\langle k|
\]
is implemented from a carry-lookahead adder [2410.05241].

A different structured extension appears for one-pair semiseparable matrices. A real symmetric one-pair semiseparable matrix \(S\) generated by vectors \(u\) and \(v\) satisfies
\[
S_{ij}=
\begin{cases}
u_i v_j,& i\le j,\\
u_j v_i,& i\ge j,
\end{cases}
\]
and admits the factorization
\[
S=D_u\,L\,\Delta_z\,L^T\,D_u,
\]
with \(D_u\) diagonal, \(L\) the lower-triangular all-ones matrix, and \(\Delta_z\) diagonal. The key subroutine is an \((\alpha_D,1,\varepsilon)\)-block-encoding of an arbitrary diagonal matrix \(D\), built from an oracle \(O_D\), a reversible Boolean circuit computing a fixed-point \(\arcsin\), and a \(U_\theta\) gadget of controlled \(R_y\)-rotations plus a controlled-\(Z\) for the sign. The full semiseparable construction uses \(2\log(N)+7\) ancillary qubits and polylogarithmic depth [2603.19130].

Taken together, these results suggest a broader structural interpretation of diagonal block-encoding: the diagonal primitive is not limited to already diagonal matrices, but also serves as the central component in factorizations or decompositions of matrices with periodic, semiseparable, or nearly circulant structure.

## 4. Circuit depth, gate counts, and synthesis optimization

Reported resource statements vary sharply with the structure being exploited.

| Construction | Reported resources |
|---|---|
| Simple diagonal \(D=C(\omega)\) | \(G_{\rm diag}(n)=\Theta(n)\) |
| Banded periodic \(M=\alpha_0C+\alpha_1L+\alpha_2R+\alpha_3I\) | \(\Theta(n^2)\) with naive shift, \(\Theta(n)\) with optimized \(O(n)\) adder |
| Finite-difference Poisson matrices | ancillas \(O(\log N)\), Toffoli-depth \(O(\log\log N)\), overall gate count \(O(\log N)\) |
| One-pair semiseparable matrices | \(2\log N+7\) ancillas, polylogarithmic depth, error \(O(N^2\max(\epsilon_u,\epsilon_v))\) |

For the Poisson matrices, the key asymptotic statement is
\[
\mathrm{Depth}(\bar L)=O(\log\log N),
\]
because the only nontrivial subroutines are multi-controlled Pauli gates of \(O(\log n)\) depth and the incrementer \(\mathrm{ADD1}\) of depth \(O(\log n)\), with \(n=\log N\). The paper further states ancilla qubits \(a=O(1)\) for the LCU plus \(O(n)=O(\log N)\) for arithmetic, overall gate count \(O(n)=O(\log N)\), and compares this against an “oracle” approach of polylog \(N\), poly \(d\) depth and a classical sparse-matrix encode of depth \(O(N)\) [2410.05241].

For periodic diagonal structure, the simple diagonal case is linear in \(n\), whereas the full banded periodic matrix depends on how shifts are implemented. PREP and PREP\(^\dagger\) on two ancillas cost \(O(1)\) two-qubit gates if implemented by small fixed unitaries. The SELECT stage is dominated either by naive increment/decrement circuits of \(O(n^2)\) Toffoli cost or by optimized \(O(n)\) adders, leading to
\[
G_{\rm band}(n)=
\begin{cases}
\Theta(n^2),& \text{naive shift},\\
\Theta(n),& \text{with optimized }O(n)\text{ adder}.
\end{cases}
\]
This separates the genuinely diagonal part from the shift overhead [2602.10589].

A separate synthesis result improves generic block-encoding circuits through “diagonal matrix migration.” In the single-ancilla protocol for \(2^{n-1}\times 2^{n-1}\) matrices, a data-qubit diagonal operator commutes through a uniformly controlled \(R_z\) layer, allowing diagonal corrections to be absorbed into subsequent stages of a block-ZXZ recursion. The resulting leading C-NOT count for general block-encoding is
\[
N_{\rm mat}(n-1)=\frac{11}{48}\,4^n+\mathcal O(2^n),
\]
and for a rank-\(K\) matrix,
\[
N_{\rm mat}^{(\mathrm{rank}\,K)}
\le
\bigl(K+\tfrac{11}{12}\bigr)2^n+\mathcal O(Kn^2).
\]
The paper also states the counting lower bound
\[
k\ge \Bigl\lceil \tfrac18\,4^n-\tfrac34\,n\Bigr\rceil
\]
for a general block-encoding of a \(2^{n-1}\times 2^{n-1}\) complex matrix [2603.16492].

The resource picture is therefore heterogeneous: diagonality can collapse costs to \(O(n)\), but neighboring structural requirements such as shifts, factor compositions, or generic state preparation may dominate the actual circuit.

## 5. Role in PDE simulation, QSVT, and parameterized scientific computing

Diagonal block-encodings are especially prominent in PDE-oriented quantum algorithms because discretized coefficient fields often appear as diagonal matrices. In the parameterized PDE framework, diagonal coefficient operators \(C_f\) are combined with finite-difference block-encodings \(D_\mu^\pm\) to obtain block-encodings of first-order and second-order PDE generators. Under a Fourier assumption for the coefficients, the resulting gate count is
\[
O\bigl(d\,K^d + d\,n\log K + n^2\bigr),
\]
with ancilla cost \(a=O(d\log K)\) for the diagonal LCU-Fourier part and error \(\varepsilon_f\) controlled by Fourier truncation [2603.01358].

Once a block-encoding \(U_A\) of \(A\) is available, QSVT is used to implement \(e^{-At}\) or another polynomial approximation on the singular values. The corresponding sequence length is
\[
L=O(\alpha_A t+\log(1/\varepsilon_{\rm sim})).
\]
The same framework extends to parameter-dependent operators
\[
\tilde A(\xi)=\sum_\xi |\xi\rangle\langle\xi|\otimes A(\xi),
\]
so that a design register can coherently control PDE evolution and objective evaluation in PDE-constrained optimization [2603.01358].

The detailed two-dimensional wave-equation example uses a shifted Gaussian speed profile
\[
c(x,y;\xi)
=
-\exp\!\Bigl(-\tfrac{(x-\xi_x)^2}{2\sigma_x^2}
-\tfrac{(y-\xi_y)^2}{2\sigma_y^2}\Bigr)+1,
\qquad
\sigma_x=1/20,\ \sigma_y=1/5,
\]
approximated by a truncated Fourier series of order \(K_x=K_y=3\). The induced coefficient operator \(\tilde C(\xi)\) is block-encoded with \((\sum|\lambda|,\log(49),0)\) parameters. After adding another 6 ancillas for the \(4\times4\) system selector and finite-difference block-encodings, the forward simulation uses 20 qubits; turning on the \((\xi_x,\xi_y)\) registers yields 28 total qubits. The reported numerical simulations reproduce wave-front slowing near the Gaussian bump and show an objective optimum near \(\xi\approx(0.5,0.5)\) [2603.01358].

The same diagonal and near-diagonal methodology appears in other PDE-related settings. For the periodic-diagonal banded matrix
\[
M=\alpha_0C(\omega,\phi)+\alpha_1L+\alpha_2R+\alpha_3I,
\]
the block-encoding is intended for QSVT tasks such as matrix inversion and matrix exponentiation, with advection-diffusion-reaction dynamics given as an application [2602.10589]. For finite-difference Poisson systems, combining the \(O(\log\log N)\)-depth block-encoding with the optimal adiabatic quantum linear solver of Costa et al. yields overall wall-clock depth
\[
O\bigl(\kappa\,\log(1/\epsilon)\,\log\log N\bigr),
\]
for preparing a quantum state representation of the PDE solution [2410.05241].

## 6. Error models, normalization, and implementation trade-offs

A central distinction in the literature is between exact and approximate diagonal block-encodings. The finite-difference Poisson constructions are exact at the block-encoding level because all \(U_j\) are exactly unitary and the LCU coefficients are uniform. By contrast, Fourier-based PDE coefficient encodings incur truncation error \(\varepsilon_f\), and semiseparable encodings accumulate approximation error from the oracle-loaded diagonal data, the fixed-point \(\arcsin\) circuit, and the product of multiple block-encodings [2410.05241], [2603.01358], [2603.19130].

Normalization can also be decisive. In the semiseparable case, the scaling factor is \(O(N^2)\), and the final error bound is
\[
\varepsilon_{\rm final}=O\bigl(N^2\max(\epsilon_u,\epsilon_v)\bigr).
\]
This means that polylogarithmic depth does not by itself imply small effective query complexity, because normalization and approximation constants still enter higher-level algorithms [2603.19130].

A separate trade-off appears in fault-tolerant implementations of diagonal quadratic operators such as
\[
U=e^{-it\phi_x^2}.
\]
In the LCU/block-encoding setting, the qubit signed-binary baseline uses projectors \(P_{r,s}\) with per-call cost
\[
T_{\rm qb}^{\rm LCU}(\varepsilon_{\rm BE})
=
32\,b_r+24\,n_b-116,
\qquad
b_r=\Bigl\lceil \tfrac12\log_2\!\bigl(9\pi^2/2\varepsilon_{\rm BE}\bigr)\Bigr\rceil,
\qquad
n_b=\lceil\log_2 d\rceil.
\]
The native qudit LCU expands \(\phi_x^2\) in generalized Pauli powers \(Z_d^r\), with PREP and SELECT requiring \(3d-3\) embedded two-level rotations in the fixed-encoding model. The paper states that the qubit encoding is asymptotically cheaper in \(d\), while finite-\(d\) threshold analysis identifies regions where qudits can offer constant-factor savings: at \(\varepsilon_{\rm sim}=10^{-6}\), \(a_{\max}>a_{R_z}\) only for \(d=3,5\) in the precision-dominated regime \(t=0.1\), and up to prime \(d\le 19\) in the time-dominated regime \(t=3000\) [2604.26792].

These results indicate that diagonal structure alone does not fix the full cost profile. The dominant overhead can lie in PREP, in shifts or adders, in normalization, in approximation error, or in fault-tolerant rotation synthesis. The common advantage of diagonal block-encodings is therefore not the elimination of all complexity, but the conversion of otherwise unstructured matrix input into a form compatible with explicit circuit analysis, QSVT composition, and structure-dependent asymptotic improvements.

Source: https://www.emergentmind.com/topics/diagonal-block-encodings