- The paper introduces PRISM, an end-to-end differentiable framework that jointly fits multi-compartment tissue models, fiber orientations, spatial regularization, and intensity calibration across voxel patches.
- PRISM recovers crossing fibers with 3.5° angular error and 95% recall in MSE mode, improving to 2.3° error and 99% recall with learned-noise Rician likelihood fitting.
- The method improves phantom tractography connectivity to r=.934 and fits a 741,000-voxel human brain in about 12 minutes on one GPU, while remaining limited by fixed diffusivities, missing distortion modeling, and limited in-vivo validation.
Overview
PRISM is a differentiable analysis-by-synthesis framework for diffusion MRI (dMRI) microstructure fitting that replaces the conventional pipeline of voxelwise, per-stage estimation with a single end-to-end optimization over spatial patches. The method fits an explicit multi-compartment forward model — combining CSF, gray matter, up to K white-matter fiber compartments (stick-and-zeppelin), and a restricted isotropic compartment — jointly with a nuisance calibration module comprising a smooth bias field, per-measurement affine intensity correction, and a learnable noise level σ. Two data-fidelity modes are supported: a fast MSE objective and a Rician negative log-likelihood (NLL) that recovers σ self-supervisedly without oracle information. The work is positioned against scalar models (DTI, DKI), orientation estimators such as CSD and MSMT-CSD, deep-learning regressors, and forward-simulation/dictionary methods (ODF-FP, FORCE); its distinguishing regime is the multi-compartment inverse problem on clinical-style acquisitions with neighbor-aware regularization and joint calibration (2604.00250).
Forward model and parameterization
The synthesized measurement at voxel x and measurement n is y^(x,n)=A(S(x,n;θmicro);θcal), where the tissue signal is a softmax-weighted mixture of mono-exponential isotropic compartments (Dcsf=3.0, Dgm=0.9, Dres=0.2×10−3mm2/s) and K stick-and-zeppelin fibers with fixed σ0, σ1, and learnable intra-axonal fraction. The restricted compartment captures slowly decaying high-σ2 signal from cellular structures in the spirit of SANDI-type models.
The calibration cascade applies an exponentiated trilinearly upsampled σ3 control-grid bias field, a per-measurement affine correction σ4, and a learned noise scale σ5. Because the product σ6 is scale-ambiguous, well-posedness rests on the frequency separation between the low-resolution bias grid and per-voxel σ7, plus σ8 penalties anchoring all calibration parameters to identity. This is an explicit assumption: on data where nuisance structure is not low-frequency or affine-per-volume, the decomposition may be ill-posed.
Regularization includes a Huber-Laplacian spatial smoothness prior over tissue fractions (6- or 26-connected; ablation shows 26-connectivity is not consistently better), direction repulsion σ9, σ0 sparsity on minor fiber fractions, orphan-WM suppression, directional continuity, and fiber ordering to break permutation symmetry. Optimization uses Rprop, selected empirically among eight optimizers for its sign-based updates that handle heterogeneous parameter scales; it attains the lowest final MSE (σ1) and fastest convergence (22 iterations to threshold) versus Adam-family baselines on Stanford HARDI slices. Convergence requires 50–300 iterations depending on dataset.
Crossing-fiber recovery
On the synthetic 16-angle benchmark (SNR = 30, 3 shells, 3400 voxels), PRISM achieves σ2 overall best-match angular error with 95% recall in MSE mode — a σ3 reduction relative to the strongest baseline, MSMT-CSD (σ4, 83% recall). In NLL mode with learned σ5, error falls to σ6 with 99% recall and F1, resolving crossings down to σ7. Notably, CSD and MSMT-CSD largely detect only one fiber at crossings σ8, while ODF-FP's dictionary discretization degrades at wider angles (σ9). Two aspects of this comparison deserve emphasis: PRISM uses a deliberately mismatched forward model (x0 vs. true x1), whereas CSD derives its response from ground-truth single-fiber voxels and MSMT-CSD receives oracle response functions — so the reported advantage is obtained under a handicap in model specification. However, the benchmark is synthetic with known ground truth, and results at SNR 10 and 50 are mentioned but not detailed in the provided content.
Under per-measurement gain perturbations (x2–x3), the calibration module halves angular error at x4 (x5) and removes 85% of reconstruction MSE, while self-regularizing to identity on clean data — confirming the module activates only under corruption rather than absorbing biological signal.
Tractography and in-vivo validation
On the DiSCo1 digital phantom (16 ROIs, 120 ROI pairs, SNR = 50), identical deterministic tracking shows PRISM (NLL mode) outperforming both CSD baselines at every tracking angle, reaching best connectivity correlation x6 at x7 versus x8 for MSMT-CSD at x9, with the largest gains (n0 pp) at tighter angles (n1–n2). The shift of PRISM's optimum to a narrower tracking angle indicates its recovered orientations support more stringent angular thresholds.
Ablation on DiSCo1 identifies the restricted compartment as the dominant contributor: removing it drops n3 from n4 to n5 (n6 pp), compared with n7 pp for the topology prior and n8 pp for the spatial prior. This is a strong claim about model composition — high-n9 restricted diffusion, not fiber-count priors, drives most of the tractography gain on this phantom.
In-vivo whole-brain fitting on HCP subject 100307 (~741k voxels, 288 measurements) completes in ~12 minutes on a single RTX A6000 in MSE mode, reaching MSE y^(x,n)=A(S(x,n;θmicro);θcal)0. Tissue fractions agree with FreeSurfer y^(x,n)=A(S(x,n;θmicro);θcal)1 segmentations at Dice 0.79 (WM) and 0.80 (GM) with Pearson y^(x,n)=A(S(x,n;θmicro);θcal)2, without any y^(x,n)=A(S(x,n;θmicro);θcal)3 input to the fit, and all summary metrics vary by less than 0.003 across three random seeds. These agreement figures are moderate rather than near-ceiling, which is expected given that dMRI-derived fractions and histological-style segmentations measure related but distinct quantities.
Limitations and open questions
The paper concedes several boundaries explicitly. PRISM models only intensity-domain nuisance factors; geometric artifacts (motion, susceptibility distortion) remain the responsibility of upstream preprocessing. Convergence relies on random initialization within a nonconvex landscape, and the authors identify warm-starting from dictionary methods (ODF-FP, FORCE) or a self-supervised encoder as untested accelerations. Fiber dispersion is not modeled (no Bingham distribution), and fixed diffusivities across all compartments impose biophysical rigidity whose sensitivity is quantified only through the single deliberate mismatch experiment. Whether the DiSCo1 ablation ranking — restricted compartment dominating — transfers to in-vivo data at clinical SNR remains open, as does scaling behavior beyond y^(x,n)=A(S(x,n;θmicro);θcal)4–y^(x,n)=A(S(x,n;θmicro);θcal)5 fibers on real acquisitions.
Conclusion
PRISM demonstrates that treating microstructure estimation as differentiable analysis-by-synthesis — with explicit fiber parameters, a restricted compartment, learned-y^(x,n)=A(S(x,n;θmicro);θcal)6 Rician fitting, and joint nuisance calibration — yields measurable improvements in crossing-fiber resolution (y^(x,n)=A(S(x,n;θmicro);θcal)7 vs. y^(x,n)=A(S(x,n;θmicro);θcal)8 angular error; y^(x,n)=A(S(x,n;θmicro);θcal)9 in NLL mode), phantom tractography agreement (Dcsf=3.00), and anatomically consistent whole-brain maps at clinically practical compute cost (~12 min per brain on one GPU). The evidence base is strongest on synthetic and phantom data with ground truth; the in-vivo validation is limited to segmentation agreement and reproducibility rather than downstream task performance, leaving quantitative in-vivo fiber-recovery accuracy as the principal open question.