Papers
Topics
Authors
Recent
Search
2000 character limit reached

Bernoulli Micro-Kernel for Quantum PDE Sampling

Updated 23 November 2025
  • Bernoulli micro-kernel is a quantum computational primitive that performs explicit stencil node updates in finite-difference PDE solvers using shallow, single-qubit circuits.
  • It leverages constant-resource Monte Carlo sampling to estimate convex-sum stencil updates, ensuring unbiased estimators with error convergence of O(1/√M).
  • Empirical evaluations on simulators and NISQ devices demonstrate its scalability, lower bias, and improved accuracy compared to deeper, entangling alternatives.

The Bernoulli micro-kernel is a quantum computational primitive designed to perform explicit stencil node updates arising in finite-difference solvers for Partial Differential Equations (PDEs). In this context, it serves as a localized, constant-resource Monte Carlo subroutine—implementable via shallow, single-qubit quantum circuits—for accelerating the sampling of convex-sum stencil updates. Its resource cost in qubits and circuit depth does not scale with the problem size, rendering it suitable for orchestrated, massively parallel applications over computational grids. The Bernoulli micro-kernel is a realization of the broader QPU micro-kernel concept, wherein a quantum processor (QPU) acts as a sampling accelerator, invoked by a classical host that maintains the outer iteration structure (Markidis et al., 16 Nov 2025).

1. Stencil Computation Framework and the Micro-Kernel Paradigm

Explicit finite-difference PDE solvers, such as the Forward-Time Centered-Space (FTCS) method for the 1D Heat equation, update the value at each spatial node ii at time n+1n+1 according to a convex combination of neighbor values at time nn:

uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.

In the QPU micro-kernel framework, the classical host iterates over time steps nn and nodes ii, invoking the quantum micro-kernel only to obtain unbiased Monte Carlo estimates of the stencil update for each node. This approach offloads the local convex-sum operation to the quantum device, leaving the global time loop and grid iteration on the classical host (Markidis et al., 16 Nov 2025).

2. Bernoulli Micro-Kernel Circuit Structure

The Bernoulli micro-kernel operates via single-qubit circuits for each stencil branch b∈{L,C,R}b\in\{L,C,R\}:

  • The qubit is initialized in ∣0⟩\ket{0}.
  • A single-qubit RyR_y rotation is applied with angle θ(ub)=2arcsin⁡ub′\theta(u_b) = 2\arcsin\sqrt{u_b'}, where n+1n+10 is the affine-normalized neighbor value mapped to n+1n+11.
  • The qubit is measured in the computational basis, yielding a Bernoulli sample with n+1n+12.
  • This process is repeated n+1n+13 times per branch to obtain an empirical mean n+1n+14.

No entanglement or multi-qubit operations are involved; each branch is executed independently (Markidis et al., 16 Nov 2025).

3. Data Encoding and Shot Allocation

Neighbor values n+1n+15 originally in n+1n+16 are linearly normalized to n+1n+17 in n+1n+18:

n+1n+19

Given a per-node shot budget nn0, shots are allocated proportionally to the stencil weights:

nn1

Each branch executes its nn2 circuit nn3 times, enabling shot-based statistical estimation respecting the convex weights (Markidis et al., 16 Nov 2025).

4. Estimator Construction and Statistical Properties

Let nn4 be the outcome of the nn5-th measurement for branch nn6. The convex-sum estimator for nn7 is

nn8

where nn9.

The estimator is unbiased:

uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.0

and its variance is

uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.1

The standard error vanishes as uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.2 (Markidis et al., 16 Nov 2025).

5. Resource Requirements and Scaling

The Bernoulli micro-kernel achieves resource independence from grid size:

  • Qubit count per branch: 1 qubit.
  • Circuit depth per shot: one uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.3 gate plus measurement.
  • No entanglement between qubits; no increase in circuit complexity with additional grid points.

This constancy makes micro-kernels amenable to large-scale grid parallelization, with classical orchestration handling all node and branch-level iteration (Markidis et al., 16 Nov 2025).

6. Error Behavior and Convergence

The standard error for each branch's mean estimator is

uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.4

and propagates through the convex-sum to

uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.5

Empirical studies using noiseless simulators confirm the uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.6 convergence: doubling uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.7 reduces the estimator noise by approximately uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.8 (Markidis et al., 16 Nov 2025).

7. Empirical Evaluation on Simulators and Quantum Hardware

Benchmarks were conducted for both the Heat and viscous Burgers’ equations:

Hardware Circuit Depth Gates Errors (uin+1=wL ui−1n+wC uin+wR ui+1n,wL+wC+wR=1,wb≥0.u_i^{n+1} = w_L\,u_{i-1}^n + w_C\,u_i^n + w_R\,u_{i+1}^n, \qquad w_L + w_C + w_R = 1, \quad w_b\ge0.9) Per-Node Wall Time
Simulator 1 1 × nn0 nn1 Not reported
IBM Brisbane 3 1 × nn2, 1 × X 0.0848, 0.0368 (raw) ≈ 4.7 s (M=4000)
0.0756, 0.0378 (mitigated)

On IBM Brisbane, the Bernoulli micro-kernel with nn3 shots per node achieved nn4, nn5 without readout mitigation and nn6, nn7 after applying single-qubit readout calibration. Circuit depth after transpilation was 3, with no two-qubit gates, and per-node wall time was approximately 4.7 s. In contrast, the branching micro-kernel exhibited higher error and deeper, more resource-intensive circuits (Markidis et al., 16 Nov 2025).

The results demonstrate that on present-day NISQ devices, the shallow, single-qubit Bernoulli micro-kernel consistently yields lower bias and higher accuracy relative to deeper, entangling alternatives, which are more susceptible to device noise (Markidis et al., 16 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Bernoulli Micro-Kernel.