---
title: 'Bucketed SPAI-GMRES-IR: Mixed-Precision Solver'
url: https://www.emergentmind.com/topics/bucketed-spai-gmres-ir
type: topic
---

# Bucketed SPAI-GMRES-IR: Mixed-Precision Solver

Bucketed SPAI-GMRES-IR is a mixed-precision Krylov subspace method for solving sparse linear systems, leveraging bucketed (adaptive-precision) sparse approximate inverse (SPAI) preconditioning within a GMRES-based iterative refinement (IR) framework. The method targets reduction in computational cost and memory consumption enabled by recent hardware trends supporting multiple floating-point precisions, while maintaining accuracy guarantees equivalent to uniform-precision approaches under specified condition number regimes [2307.03914] [2303.04251].

## 1. Construction of the Bucketed SPAI Preconditioner

Let $A \in \mathbb{R}^{n\times n}$ be a nonsingular matrix, and let $M \approx A^{-1}$ denote a right preconditioner computed using the Frobenius-norm SPAI algorithm. Once $M$ is constructed (typically in higher precision), its nonzero entries are partitioned into “buckets” according to their magnitudes, with each bucket assigned a corresponding precision level.

Given $q \geq 2$ decreasing unit roundoffs $u_1 > u_2 > \ldots > u_q$ and a target threshold $\epsilon_B \leq \min(u_1,\ldots,u_q)$, define for each row $i$ and for $k=1,\ldots,q$ the intervals
- $P_{i1} = (\epsilon_B\|M\|/u_2,\,+\infty)$,
- $P_{ik} = (\epsilon_B\|M\|/u_{k+1},\ \epsilon_B\|M\|/u_k]$, $k=2,\ldots,q-1$,
- $P_{iq} = [0,\ \epsilon_B\|M\|/u_q]$.

The bucket $B_{ik}$ for row $i$ and precision $u_k$ consists of column indices $j$ such that $|m_{ij}| \in P_{ik}$. After bucketing, a SpMV $y = Mx$ is evaluated by computing, for each $i$ and $k$, $y_i^{(k)} = \sum_{j \in B_{ik}} m_{ij} x_j$ in precision $u_k$, then summing $\sum_{k=1}^q y_i^{(k)}$ in the highest precision $u_1$ [2307.03914].

## 2. Mixed-Precision GMRES-IR Framework with Bucketed SPAI

Bucketed SPAI-GMRES-IR embeds this adaptive-precision preconditioner into a five-precision GMRES-based iterative refinement. The relevant precisions are:

- $u_f$ for preconditioner construction,
- $u_r$ for residual computation,
- $u$ for working storage,
- $u_g$ for GMRES arithmetic,
- $u_p$ for $A$-vector products.

The preconditioner apply ($M r_i$) in each GMRES step uses bucketed SpMV as above. All other core Krylov and orthogonalization operations remain in high precision to ensure algorithmic stability. The iterative refinement proceeds as follows:

1. Compute $M$ (SPAI) in precision $u_f$.
2. Bucket $M$ entries according to the adaptive-precision rules.
3. Compute initial solution $\hat x_0 = M b$.
4. For each of $L$ outer iterations:
   - Compute residual $r_i = b - A x_i$ in $u_r$.
   - Use left-preconditioned GMRES to solve $\tilde{A} d_i = M r_i$ to tolerance $\tau$, where $\tilde{A} = M A$.
   - Update $x_{i+1} = x_i + d_i$ in precision $u$.

Backward and forward error checks determine convergence [2307.03914] [2303.04251].

## 3. Convergence Guarantees and Stability Analysis

The bucketed SpMV induces a perturbed preconditioner $M + \Delta M$, with $\|\Delta M\|_F \leq ((q-1)u_1 + c\epsilon_B)\|M\|_F$ where $c$ depends mildly on the partition sizes. If $\epsilon_B \approx u_1 \approx u_p$, the additional error from bucketed evaluation does not degrade convergence relative to uniform-precision SPAI-GMRES-IR. Specifically, for suitably chosen precisions and bucket threshold, GMRES with bucketed $M$ satisfies

$$
\| \widehat{y} - y \| / \|y\| \leq (q-1)u_1 + c \epsilon_B,
$$

and the GMRES-IR outer iteration converges with backward/forward error $O(u)$ under the same spectral conditions as uniform-precision preconditioning [2307.03914].

Adopting the essential-forward-and-backward stability (EFBS) paradigm, the bucketed SPAI-GMRES-IR attains forward and backward error bounds in practical scenarios that are independent of $\kappa(A)$, provided the underlying problem is well-posed and residuals are computed in high precision [2303.04251].

## 4. Computational Cost, Storage, and Precision Allocation

Memory cost for the bucketed preconditioner is

$$
\mathrm{Storage}(M)_{\text{bucketed}} = \sum_{k=1}^q p_k\, \beta(u_k),
$$

with $p_k$ the total number of entries stored in precision $u_k$ and $\beta(u_k)$ the storage cost per entry. For the uniform-precision case:

$$
\mathrm{Storage}(M)_{\text{uniform}} = \mathrm{nnz}(M) \, \beta(u_f).
$$

The storage reduction ratio,

$$
\rho_{\text{store}} = \frac{\sum_k p_k\,\beta(u_k)}{\mathrm{nnz}(M)\,\beta(u_f)} \leq 1,
$$

quantifies memory gain. Application of $M$ to a vector costs

$$
\mathrm{Cost}(\mathrm{SpMV}_B) = \sum_{k=1}^q p_k\,\gamma(u_k) + \mathrm{nnz}(M)\,\gamma(u_1),
$$

where $\gamma(u_k)$ is the per-operation compute cost at precision $u_k$. Typically $\gamma(u_k) \ll \gamma(u_f)$ for $k > 1$, enabling significant runtime and energy savings. Parameter selection (number of buckets, thresholds, precision levels) is guided by the preconditioner’s spectrum, hardware capabilities, and the target error budget [2307.03914] [2303.04251].

## 5. Numerical Results and Empirical Trade-offs

Extensive experiments on SuiteSparse matrices and synthetic ill-conditioned systems demonstrate key trade-offs:

| Matrix   | Method, $\epsilon_B$  | $\kappa_\infty(MA)$ | nnz Buckets                          | $\rho_{\text{store}}$ | GMRES Iters (per-refine) |
|----------|----------------------|---------------------|--------------------------------------|-----------------------|--------------------------|
| steam1   | SPAI $\epsilon_{SP}$ | 1.5                 | 1105 (1105,0,0,0)                    | 1.00                  | 14 (7,7)                 |
| steam1   | BSPAI $2^{-53}$      |                     | 1105 (556,537,12,0)                  | 0.749                 | 21 (7,7,7)               |
| steam1   | BSPAI $2^{-37}$      |                     | 1105 (242,284,347,232)               | 0.426                 | 21 (7,7,7)               |

Highlights:

- Where uniform SPAI-GMRES-IR converges, bucketed SPAI-GMRES-IR also converges within comparable iterations for small $\epsilon_B$.
- Storage for the preconditioner can be reduced by up to 60% with only a mild increase in GMRES iterations.
- For thresholds $\epsilon_B \approx u$, iteration count remains almost unchanged, with moderate storage reduction.

For randsvd and real-world matrices, the method achieves backward and forward errors on the order of $10^{-16}$. Preconditioner application costs are substantially reduced, and wall-time is 2–5× lower than double precision direct solves, with energy savings accruing from the cheap low-precision matvecs [2307.03914] [2303.04251].

## 6. Implementation and Practical Recommendations

Effective implementation of bucketed SPAI-GMRES-IR requires:

- Contiguous storage layouts for each bucket, enabling efficient dispatch to SIMD or GPU kernels for each precision.
- Specialized mixed-precision BLAS kernels for bucketed SpMV in half, single, and double precisions.
- Orthogonalization in high precision (e.g., classical Gram–Schmidt in double) to maintain stability.
- Parallelism is maximized by the row-wise independence in bucketed SPAI matvecs and by fusing bucketed computations to minimize synchronization overhead.
- For communication-avoiding variants in distributed-memory contexts, fusing global reductions in high precision is sufficient.

This architecture integrates readily into MPI+OpenMP or CUDA libraries, yielding EFBS-certified sparse solvers suitable for large-scale, mixed-precision HPC deployments [2303.04251].

## 7. Theoretical and Practical Significance

Bucketed SPAI-GMRES-IR enables new forms of adaptive-precision preconditioning with provable error guarantees matching those of standard uniform-precision methods, contingent on regime-appropriate parameter selection. Empirical and theoretical analyses indicate that, under well-posedness and with proper bucket thresholding, the method is robust to the degradation often associated with low-precision arithmetic, facilitating energy and cost savings without sacrificing solution quality. *A plausible implication is that further hardware trends toward mixed-precision support will amplify these gains for large-scale, sparse scientific computing* [2307.03914] [2303.04251].

Source: https://www.emergentmind.com/topics/bucketed-spai-gmres-ir