MP-LBFGS: Efficient Optimization for FBPINNs
- The algorithm leverages local LBFGS steps within FBPINN subdomains to form quasi-Newton corrections, significantly reducing global synchronizations.
- It aggregates local corrections through a nonlinear subspace minimization, preserving curvature benefits while ensuring global secant consistency.
- Empirical benchmarks on PDE problems show reduced training epochs and improved accuracy by trading off local computation for fewer communications.
Multi-Preconditioned LBFGS (MP-LBFGS) is an optimization algorithm developed to accelerate and robustify the training of finite-basis physics-informed neural networks (FBPINNs). Inspired by the nonlinear additive Schwarz method from domain decomposition in numerical PDEs, MP-LBFGS exploits the intrinsic additive structure of FBPINNs by constructing parallel, subdomain-local quasi-Newton corrections and then optimally combining them via a nonlinear, low-dimensional subspace minimization. This approach yields a preconditioned search direction that both improves convergence and reduces communication overhead in distributed learning settings (Salvadó-Benasco et al., 13 Jan 2026).
1. FBPINN Architecture and Domain Decomposition
The motivating problem is the solution of a partial differential equation (PDE) on a bounded domain with boundary conditions on . A physics-informed neural network (PINN) approximates the unknown by a global neural network , trained to minimize the squared PDE residual over collocation points.
FBPINNs address limitations such as spectral bias by decomposing into overlapping subdomains (with overlap width ). Each 0 receives a local network 1; the global solution is assembled via a smooth partition of unity 2 satisfying 3 and 4: 5 The normalization 6 maps local coordinates to 7. Collocation data and model parameters thus admit natural subdomain-local partitioning, facilitating parallelism.
2. Nonlinear Additive Schwarz Motivation
In classical PDE solvers, nonlinear additive Schwarz preconditioning accelerates Newton/quasi-Newton methods by performing local solves on each subdomain and then aggregating the results. MP-LBFGS transposes this methodology: each subnetwork solves a local training subproblem using several LBFGS steps, generating a local correction. Aggregation is enforced by a right-preconditioned LBFGS step that ensures global secant consistency.
The full-space optimality condition 8 is reformulated as
9
with the nonlinear additive Schwarz map
0
where 1 restricts global parameters to subdomain 2, and 3 approximates the local minimum via local LBFGS updates. This induces a lifted, preconditioned variable on which global LBFGS is performed, retaining curvature benefits while reducing communication frequency.
3. Algorithmic Components of MP-LBFGS
3.1 Local LBFGS Corrections
At each global iteration 4, the global parameters 5 are distributed to each subdomain, yielding 6. Each local network minimizes its loss 7 by performing 8 LBFGS steps: 9 0, resulting in an approximate minimizer 1 and local correction 2.
3.2 Subspace Aggregation
Local corrections are collected as
3
Coefficients 4 are sought such that 5 minimizes the overall objective. While linear preconditioning would suggest
6
robustness in the neural setting is enhanced by instead solving the nonlinear subspace minimization
7
This problem is efficiently solved via a few (damped) Newton steps; the Hessian 8 is only 9 and thus computationally cheap.
3.3 Global LBFGS Step
The updated parameter is 0, with new gradient 1. The global secant pair 2 is then formed: 3 The global LBFGS memory is updated, the descent direction 4 is computed with standard recursion, and a Wolfe-condition line search yields
5
4. Iterative Structure and Communication Patterns
The MP-LBFGS outer iteration proceeds as follows:
| Step | Operation | Communication Pattern |
|---|---|---|
| 1 | Compute 6 | Global all-reduce |
| 2 | Parallel local LBFGS | None; fully local |
| 3 | Aggregate local corrections | Minimal (local to global) |
| 4 | Subspace minimization | Fully local or negligible |
| 5-10 | Global LBFGS/line search | Standard (as in vanilla LBFGS) |
Relative to traditional LBFGS, MP-LBFGS reduces the number of global synchronizations by a factor 7, replacing some forward/backward passes over the full network with local computations.
5. Computational Complexity per Outer Iteration
Let 8 be the total parameter count, 9 per subdomain, 0 the global LBFGS memory size, and 1 the number of local steps:
- Forward/backward passes: 2 global gradient eval plus 3 extra loss evals; 4 local gradients per subdomain in parallel. Subspace minimization involves a few parallel evaluations.
- LBFGS recursion: 5 flops and 6 memory for global step; 7 per local memory.
- Communication: One all-reduce per global gradient and per line search; MP-LBFGS thus reduces the number of synchronizations by 8 compared to standard LBFGS.
This configuration enables a tradeoff between local computational work and synchronization frequency.
6. Empirical Performance and Benchmarks
Numerical experiments on 1D and 2D Poisson equations and the 2D time-dependent Burgers' equation used FBPINN configurations with four-layer ResNets (width 20), 2,000–20,000 Hammersley-sampled collocation points, and various uniform decompositions (9, 0, 1).
Scaling strategies compared:
- Uniform scaling (UniS)
- Line-search scaling (LSS)
- Subspace minimization (SPM)
Key empirical findings:
- SPM was most robust and stable as 2 increased.
- MP-LBFGS with SPM reduced the required outer epochs (global synchronizations) by up to an order of magnitude compared with standard LBFGS.
- Total per-device gradient work was comparable or lower; final validation error in 3 norm was often an order of magnitude smaller.
- Increasing local LBFGS steps 4 further reduced synchronizations, trading increased local computation.
7. Practical Recommendations and Insights
The nonlinear Schwarz-type preconditioner enabled by the FBPINN decomposition allows extensive local computation, substantially reducing communication overhead in distributed settings. Subspace minimization is critical for stability: uniform or sequential line searches fail to scale as the number of subdomains increases.
Implementation recommendations:
- Tune 5 (local steps), 6 (memory size), and 7 (number of subdomains) to balance local computation against synchronization.
- Retain separate local and global LBFGS memories to avoid mixing curvature information.
- Precompute and cache small Hessian blocks 8 to expedite subspace Newton solves.
- MP-LBFGS can be integrated into existing FBPINN codebases by wrapping local training in an outer loop, collecting corrections, solving the 9-dimensional subspace problem, then applying a standard global LBFGS update.
In summary, MP-LBFGS leverages the additive decomposition of FBPINNs to perform parallel, local quasi-Newton updates, followed by a global preconditioned update whose direction is defined by a small, nonlinear subspace optimization. This mechanism accelerates convergence, lowers wall-clock time in distributed environments, and can improve final model accuracy relative to standard LBFGS (Salvadó-Benasco et al., 13 Jan 2026).