---
title: Diagonal-Parallel Picard in Fixed-Point Methods
url: https://www.emergentmind.com/topics/diagonal-parallel-picard
type: topic
---

# Diagonal-Parallel Picard in Fixed-Point Methods

Searching arXiv for papers directly relevant to “Diagonal-Parallel Picard” and adjacent fixed-point/parallel-Picard formulations.
Diagonal-Parallel Picard is not a standardized method name in the recent arXiv literature. The exact phrase does not appear in the principal papers most closely associated with parallel Picard-style methods. Instead, the phrase is best read as a compact label for several closely related constructions: Picard-style fixed-point iterations made parallel by a diagonal approximation, a diagonal preconditioner, or a coordinatewise decoupling that leaves only a scan-, map-, or component-parallel subproblem. In the literature on parallelizing sequential models, the closest explicit meaning is a diagonal-Jacobian quasi-Newton update inside a unifying linear dynamical systems framework; in spectral deferred corrections and collocation, it corresponds more literally to diagonal-preconditioned Picard iteration across collocation nodes; and in multidimensional reflected \(G\)-BSDEs it denotes Picard iteration under diagonal generators, where each frozen-coordinate update is a one-dimensional reflected equation [2509.21716] [1703.08079] [2108.09012].

## 1. Terminological status and scope

Several recent papers explicitly state that the exact phrase “Diagonal-Parallel Picard” does not appear in their text. In the parallel-sequential-model literature, the nearest named objects are Picard iteration with identity transition, quasi-Newton iteration with diagonal Jacobian approximation, and more generally parallel fixed-point methods obtained by choosing structured approximations of the Jacobian that remain closed under composition. In diffusion and Metropolis sampling, the nearest objects are global Picard updates over an entire trajectory, sometimes with a Jacobi-like time-parallel interpretation or with a moving online frontier rather than a diagonal matrix structure [2509.21716] [2503.19731] [2506.09762].

This suggests that “Diagonal-Parallel Picard” functions less as a canonical algorithm name than as an umbrella description spanning at least three meanings of “diagonal.” In one meaning, “diagonal” refers to a diagonal Jacobian approximation in state space. In another, it refers to a diagonal preconditioner in collocation-node space. In a third, it refers only informally to a diagonal or frontier-like propagation pattern through the time-index/iteration grid. The common invariant across these uses is a Picard-style fixed-point backbone combined with a structural simplification that exposes parallelism.

## 2. Parallel fixed-point iteration for sequential models

A central recent formulation studies the sequential nonlinear recursion
\[
x_{t+1}=f_{t+1}(x_t),
\]
recasts rollout as a fixed-point problem over the full trajectory, and shows that Newton, quasi-Newton, Picard, and Jacobi all fit the same linear-dynamical-systems update
\[
x_{t+1}^{(i+1)}=f_{t+1}(x_t^{(i)})+A_{t+1}\bigl(x_t^{(i+1)}-x_t^{(i)}\bigr).
\]
The method is determined by the choice of \(A_{t+1}\): full Newton uses the full Jacobian, quasi-Newton uses the diagonal Jacobian, Picard uses \(I_D\), and Jacobi uses \(0\). The paper states explicitly that “Picard iterations are in fact a type of quasi-Newton iteration, where we approximate the Jacobian of the dynamics function by the identity matrix,” while the diagonal approximation is
\[
A_{t+1}=\operatorname{diag}\!\left[\frac{\partial f_{t+1}}{\partial x_t}(x_t^{(i)})\right].
\]
Within this framework, the strict paper-grounded interpretation of “Diagonal-Parallel Picard” is therefore not Picard in the paper’s own terminology, but the diagonal quasi-Newton member of the same family [2509.21716].

The same paper makes the computational reason precise. Once \(A_{1:T}\) and the previous iterate are fixed, the update becomes an affine linear time-varying recurrence over \(t\), and affine-map composition is associative. The full trajectory can therefore be evaluated by a parallel scan in \(\mathcal{O}(\log T)\) depth on \(\mathcal{O}(T)\) processors. The diagonal structure matters because diagonal matrices are closed under multiplication, giving a cheap scan representation. The paper records the cost gap explicitly: full Newton requires \(\mathcal{O}(T D^2)\) memory and \(\mathcal{O}(T D^3)\) work, whereas diagonal quasi-Newton requires \(\mathcal{O}(T D)\) space and \(\mathcal{O}(T D)\) work [2509.21716].

A later thesis places the same family inside a broader parallel Newton framework. There the generic update is
\[
s_t^{(i+1)}=f_t(s_{t-1}^{(i)})+\tilde A_t\bigl(s_{t-1}^{(i+1)}-s_{t-1}^{(i)}\bigr),
\]
with \(\tilde A_t=A_t\) for Newton, \(\tilde A_t=\operatorname{diag}(A_t)\) for diagonal quasi-DEER, \(\tilde A_t=I_D\) for Picard, and \(\tilde A_t=0\) for Jacobi. The thesis is explicit that the phrase “Diagonal-Parallel Picard” is not defined there; the closest explicit method is diagonal quasi-DEER, while Picard itself remains the identity approximation [2603.16850].

## 3. Diagonal-preconditioned Picard in collocation and spectral deferred correction

In collocation-based time integration, Picard iteration arises from the collocation system
\[
\left(\mathbf{I}-\Delta t\,Q\mathbf{F}\right)(\mathbf{u})=\mathbf{u}_0,
\]
and spectral deferred correction is written as a preconditioned Picard iteration
\[
\left(\mathbf{I}-\Delta t\,Q_\Delta \mathbf{F}\right)(\mathbf{u}^{k+1})
=
\mathbf{u}_0+\Delta t\,(Q-Q_\Delta)\mathbf{F}(\mathbf{u}^k).
\]
Classical SDC uses a lower-triangular \(Q_\Delta\), which makes the sweep Gauss-Seidel-like and serial across nodes. The diagonal-parallel variant replaces that by a diagonal \(Q_\Delta\), so that all node updates decouple and can be computed simultaneously. One paper tests three such diagonal preconditioners:
\[
Q_\Delta^{\mathrm{Qpar}}=\operatorname{diag}(q_{11},\dots,q_{MM}),
\]
\[
Q_\Delta^{\mathrm{IEpar}}=\operatorname{diag}(\tau_1-t_0,\dots,\tau_M-t_0),
\]
and
\[
Q_\Delta^{\mathrm{MIN}}=\operatorname{diag}(\hat{\mathbf q}),\qquad
\hat{\mathbf q}=\operatorname*{argmin}_{\mathbf q\in\mathbb{R}^M}\rho\!\left(\mathbf I-\operatorname{diag}(\mathbf q)Q\right).
\]
This is the most literal sense in which the literature contains a diagonal-parallel Picard construction: a Picard or SDC sweep whose preconditioner is diagonal in node space [1703.08079].

A later paper makes the relationship sharper by stating that plain Picard iteration is recovered by choosing \(Q_\Delta=0\), while diagonal SDC is a diagonal-preconditioned Picard iteration. It proposes analytically designed diagonal sweepers, including
\[
Q_{\Delta,M}=\operatorname{diag}\!\left(\frac{\tau_1}{M},\dots,\frac{\tau_M}{M}\right),
\]
for which
\[
(Q-Q_{\Delta,M})^M=0,
\]
and a nonstationary family
\[
Q_\Delta^{(k)}=\operatorname{diag}\!\left(\frac{\tau_1}{k},\dots,\frac{\tau_M}{k}\right),
\]
for which the stiff-limit propagator product vanishes after \(M\) sweeps. In that paper’s formulation, diagonal-parallel SDC is “Picard iteration for collocation” plus diagonal implicit preconditioning, rather than Picard in the strict unpreconditioned sense [2403.18641].

## 4. Time-parallel trajectory Picard in diffusion and Metropolis sampling

A separate line of work applies Picard iteration not to state-space Jacobians or collocation matrices, but to an entire denoising or Markov trajectory. For diffusion sampling, the trajectory-form Picard update is
\[
X^{k+1}\leftarrow \Phi(X^k),\qquad
x_t^{k+1}\leftarrow x_0^k+\frac{1}{T}\sum_{i=0}^{t-1}s_\theta(x_i^k,i/T).
\]
This is parallel across denoising time because all timestep states are updated from the previous global iterate. The paper explicitly says that it does not introduce a new diagonal block decomposition; if mapped to “diagonal-parallel” language, the closest analogy is a Jacobi-like global update over the time dimension. Its main contribution is the Picard Consistency Model, trained to predict the fixed-point trajectory more directly, together with model switching to recover exact convergence to the base sampler. The reported headline gains are “up to a 2.71x speedup over sequential sampling” and “a 1.77x speedup over Picard iteration” [2503.19731].

For zeroth-order Metropolis chains, the relevant object is the Online Picard algorithm. The chain
\[
X_{i+1}=X_i+f(X_i,W_i)
\]
is rewritten via the Picard map
\[
X_i'=\Phi_i(X,W)=
\begin{cases}
X_0,& i=0,\\
X_0+\sum_{\ell=0}^{i-1}f(X_\ell,W_\ell),&0<i\le K.
\end{cases}
\]
The online scheme advances the largest already-correct prefix and reallocates processors forward, producing what the paper describes as a moving frontier through the iteration-versus-time plane. The exact phrase “Diagonal-Parallel Picard” is not used, but this is the closest object in the paper to a diagonal-front propagation interpretation. The exact algorithm gives \(\mathcal{O}(\sqrt d)\) parallel iterations with \(\mathcal{O}(\sqrt d)\) processors for Random Walk Metropolis under the stated assumptions, while an approximate variant is claimed in the abstract to generate samples from an approximate measure \(\pi_\epsilon\) in \(\mathcal{O}(1)\) parallel iterations and \(\mathcal{O}(d)\) processors; the body adds that a rigorous analysis of the invariant-measure error is outside the scope of the work [2506.09762].

## 5. Diagonal generators and coordinate-parallel Picard for reflected \(G\)-BSDEs

In multidimensional reflected \(G\)-BSDEs, “diagonal” has a different but very precise meaning. The \(i\)-th generators are
\[
f^i(t,\omega,y,z^i),\qquad g^i(t,\omega,y,z^i),
\]
so each component may depend on the full vector \(y\in\mathbb{R}^k\), but in the \(z\)-variable depends only on the scalar \(z^i\). The reflected equation is
\[
Y_t^i=\xi^i+\int_t^T f^i(s,Y_s,Z_s^i)\,ds+\int_t^T g^i(s,Y_s,Z_s^i)\,d\langle B\rangle_s-\int_t^T Z_s^i\,dB_s+(A_T^i-A_t^i),
\]
with obstacle \(Y_t^i\ge S_t^i\) and the Skorohod-type condition that
\[
\left\{-\int_0^t (Y_s^i-S_s^i)\,dA_s^i\right\}
\]
is a non-increasing \(G\)-martingale. The Picard method freezes the cross-coordinate \(Y\)-dependence through a process \(U\), defines
\[
f^{i,U}(s,y^i,z^i)=
f^i\bigl(s,U_s^1,\dots,U_s^{i-1},y^i,U_s^{i+1},\dots,U_s^k,z^i\bigr),
\]
and then solves \(k\) one-dimensional reflected \(G\)-BSDEs independently on a short interval. The associated map \(U\mapsto Y^U\) is a contraction in \(M_G^\beta\) for sufficiently small horizon \(h\), and global existence follows by backward patching [2108.09012].

This is arguably the cleanest instance of a genuinely coordinate-parallel Picard method in the supplied literature. The diagonal assumption on the generators is not merely a convenience. It is the structural reason that each frozen-coordinate update remains a one-dimensional reflected \(G\)-BSDE, so the established one-dimensional theory applies componentwise. Without diagonal \(z\)-dependence, the per-coordinate reflected solve would not decouple.

## 6. Convergence regimes, trade-offs, and limitations

In the scan-based sequential-model framework, Newton, diagonal quasi-Newton, Picard, and Jacobi all share a finite-step propagation property: in the sequence-evaluation setting they are guaranteed to converge in at most \(T\) iterations. The practical number of iterations, however, depends on how well the approximation \(A_t\) or \(\tilde A_t\) matches the true Jacobian and on whether the induced linearized recurrence is stable. The resulting regime picture is explicit: Picard works well when the Jacobian is close to \(I_D\), diagonal quasi-Newton works well when state coordinates are nearly uncoupled, and full Newton works best when off-diagonal coupling is important [2509.21716] [2603.16850].

The same sources emphasize that cheaper structure can increase iteration count and can also create numerical instability. One paper notes that “LDS matrices with spectral norm close to or greater than one can cause numerical instabilities in the parallel scan operation.” The thesis sharpens the theoretical side by giving a convergence-rate decomposition in terms of approximation accuracy and stability, and its abstract states that “the sign of the Largest Lyapunov Exponent of a dynamical system determines whether or not parallel Newton methods converge quickly” [2509.21716] [2603.16850].

For collocation, the trade-off is analogous. Diagonal SDC or diagonalized collocation exposes node-level concurrency, but usually gives weaker preconditioning than serial LU-like sweeps. One paper reports that diagonal choices can be nearly as good in non-stiff regimes, while for stiffer problems the diagonal-minimizing choice is usually the strongest among diagonal methods and the LU-based choice is usually best overall. Another shows that carefully optimized diagonal sweepers can recover stability domains and convergence behavior close to serial SDC while preserving stage parallelism [1703.08079] [2403.18641].

## 7. Distinction from unrelated uses of “diagonal” and “Picard”

The phrase should not be confused with unrelated subjects that happen to contain the same words. In combinatorics, “parallel diagonals” refers to diagonals of polygons parallel to a fixed edge, as in the enumeration of triangulations by number of parallel diagonals [1208.3915]. In algebraic geometry, “Picard rank” or “Picard number” refers to the rank of a divisor class group or Néron–Severi group, as in work on diagonal quartic surfaces with Picard number \(1\) and infinitely many rational points [1402.4583]. In toric geometry, “diagonal” may refer to the diagonal sheaf on \(X\times X\) and “Picard rank \(2\)” to the rank of \(\operatorname{Pic}(X)\), again with no algorithmic relation to parallel Picard iteration [2208.00562].

Within numerical analysis and machine learning, by contrast, the most defensible encyclopedic definition is narrower: a Diagonal-Parallel Picard method is a Picard-style fixed-point scheme whose parallelism is exposed by a diagonal or structurally separable approximation. Depending on context, that diagonal structure may sit in a Jacobian, a collocation preconditioner, a generator, or only an informal iteration-versus-time propagation pattern. The literature does not presently standardize the phrase, but it repeatedly standardizes the underlying ingredients.

Source: https://www.emergentmind.com/topics/diagonal-parallel-picard