---
title: 'CellINR: Implicit Reconstruction in Live Cell Imaging'
url: https://www.emergentmind.com/topics/cellinr-framework
type: topic
---

# CellINR: Implicit Reconstruction in Live Cell Imaging

Searching arXiv for the CellINR paper and closely related microscopy-INR work to ground the article in current literature.
CellINR is a self-supervised, case-specific reconstruction framework for removing photo-induced artifacts from 4D live fluorescence microscopy while preserving true cellular structure. It is presented as an implicit neural representation-based approach that treats a 3D cell volume as a continuous coordinate-to-intensity function rather than as a discrete voxel grid, with the explicit aim of overcoming photobleaching- and phototoxicity-induced degradation that compromises image continuity and detail recovery during prolonged live imaging [2508.19300]. The framework combines positional encoding, blind convolution, structural amplification, and a hybrid reconstruction-plus-regularization objective to separate genuine biological signals from artifacts in a setting where pixel-aligned clean ground truth is often unavailable [2508.19300].

## 1. Problem setting and motivation

CellINR is motivated by a central limitation of long-term 4D live fluorescence microscopy: the same specimen is repeatedly scanned across time, and each 3D volume is formed by layer-by-layer optical sectioning under prolonged, high-intensity illumination [2508.19300]. The reported consequences are photobleaching, which causes fluorophores to decay and the signal-to-noise ratio to fall over time, and phototoxicity, which can damage cells and alter their behavior [2508.19300].

The framework distinguishes these degradations from ordinary random noise. According to the paper, photo-induced artifacts are spatially structured and temporally systematic: local attenuation is non-uniform, drift can accumulate over time, and pseudo-signals or background structures may appear [2508.19300]. This makes the restoration problem harder than conventional denoising, because methods that suppress high-frequency noise may fail to remove low-frequency artifact patterns, while some approaches may also destroy genuine biological continuity [2508.19300].

A further difficulty is the scarcity of pixel-aligned clean targets in microscopy datasets, which limits the applicability of standard supervised learning [2508.19300]. CellINR addresses this by optimizing each cell volume individually in a self-supervised manner. This suggests that the framework is designed less as a generic feed-forward denoiser and more as a reconstruction method that exploits the internal structure of a single acquisition [2508.19300].

## 2. Implicit representation and continuous volume modeling

The representational core of CellINR is an implicit neural representation (INR), defined as a neural network
$$
F_\theta: \mathbf{x} \mapsto F_\theta(\mathbf{x}),
$$
which maps coordinates to a desired quantity [2508.19300]. In the fluorescence-imaging setting, the network learns a mapping from spatial coordinates to pixel or color values [2508.19300].

CellINR uses an MLP with parameters \(\varrho\) and positional encoding to represent the fluorescence volume continuously:
$$
\gamma(\rho)=[\sin(\rho),\cos(\rho),\ldots,\sin(2^{\epsilon-1}\rho),\cos(2^{\epsilon-1}\rho)]^T,
$$
$$
c=F_\varrho(\gamma(\rho)),
$$
where \(\rho\) is a 3D coordinate and \(c\) is the predicted intensity or color at that point [2508.19300]. The positional encoding is used to map 3D coordinates into a high-frequency domain, because raw coordinates are low-frequency inputs and can make the network too smooth [2508.19300]. The paper states that this multiscale frequency injection enables representation of fine details, edges, and thin structures [2508.19300].

A plain INR, however, is described as insufficient for the target task. The paper notes that an unaugmented INR tends to behave like an identity mapping and is not enough to separate clean structure from artifacts [2508.19300]. CellINR therefore extends the INR backbone with blind convolution and structural amplification. In functional terms, the framework uses the continuity and coordinate-based expressivity of INRs, but supplements them with mechanisms intended to prevent trivial copying and to bias optimization toward biologically meaningful structures [2508.19300].

## 3. Blind convolution and neighborhood-based inference

Blind convolution is introduced to make the model learn from surrounding context while avoiding direct access to the target pixel, in the spirit of blind-spot self-supervision [2508.19300]. The paper starts from a traditional 3D convolution:
$$
I_{(x,y,z)} = \sum_{i=-h}^{h} \sum_{j=-h}^{h} \sum_{k=-h}^{h} C_{x-i,y-j,z-k} \cdot W_{x-i,y-j,z-k},
$$
where \(C\) denotes sampled voxel values around the center and \(W\) denotes convolution weights [2508.19300]. The stated goal is to estimate the center value without feeding the center value itself into the network, thereby preventing trivial copying [2508.19300].

To adapt this idea to INR while preserving continuity, CellINR performs coarse sampling in a cube around the target point, predicts volume density with a coarse MLP \(F_\Theta\), and then performs importance sampling to obtain finer samples:
$$
\sigma = F_\Theta(\gamma(N_c)),
$$
$$
N_f=\xi(N_c,\sigma),
$$
$$
N_{(x,y,z)}=N_c\cup N_f.
$$
The fine MLP \(F_\delta\) then predicts color values:
$$
c(\rho) = F_{\delta} \left[\gamma\left(\rho \cup \xi\left(\rho, F_{\Theta} \left[\gamma\left(\rho\right)\right]\right)\right)\right], \quad \text{where } \rho \in N.
$$
A separate MLP \(F_\phi\) generates spatially varying kernel weights that are normalized with Softmax:
$$
w(\rho) = \frac{\exp(F_{\phi} (\gamma(\rho)))}{\sum_{\rho \in N} \exp(F_{\phi} (\gamma(\rho)))}.
$$
The final blind-convolution output is
$$
I_{\text{clean} = \sum_{\rho \in N} \left[ c(\rho) \cdot w(\rho) \right].
$$
These components together allow the framework to infer the center value from nearby context while maintaining spatial continuity of cellular structures such as membranes and nuclei [2508.19300].

In implementation, within a cube of radius 1 around the blind spot, \(N=27\) points are sampled, then another \(N\) fine points are added, for a total of \(2N\) inputs to the fine MLP and kernel module [2508.19300]. At inference time, the fine network predicts intensities from sampled center coordinates and the kernel weights are set to 1, allowing output generation at arbitrary resolutions by adjusting the sampling density [2508.19300]. A plausible implication is that CellINR uses the INR formalism not only for denoising-like restoration but also for continuous resampling.

## 4. Structural amplification and loss design

The second major component is structural amplification, which is specifically intended to suppress photo-induced artifacts that are described as low-frequency, smooth, and locally uniform [2508.19300]. The paper contrasts these artifacts with true biological structures, which are said to be concentrated and high-frequency [2508.19300]. To exploit this distinction, CellINR uses a Hessian-based enhancement step. For a 3D fluorescence image \(I^f(x,y,z)\), the Hessian matrix is
$$
H = \begin{bmatrix} \frac{\partial^2 I^f}{\partial x^2} & \frac{\partial^2 I^f}{\partial x \partial y} & \frac{\partial^2 I^f}{\partial x \partial z} \\
\frac{\partial^2 I^f}{\partial x \partial y} & \frac{\partial^2 I^f}{\partial y^2} & \frac{\partial^2 I^f}{\partial y \partial z} \\
\frac{\partial^2 I^f}{\partial x \partial z} & \frac{\partial^2 I^f}{\partial y \partial z} & \frac{\partial^2 I^f}{\partial z^2}
\end{bmatrix} = E\Lambda E^T.
$$
After eigenvalue decomposition, the largest-magnitude eigenvalue \(\lambda_3\) is used to produce an enhanced image:
$$
I_{\text{en}(x,y,z)=\frac{|\lambda_3(x,y,z)|}{\max\{|\lambda_3(x,y,z)| \mid x,y,z \in I\}.
$$
Thresholding via Otsu’s method then yields a binary mask:
$$
I_{\text{binary}(x,y,z)= \begin{cases} 255, & \text{if } I_{\text{en}(x,y,z) > \mu, \\
0, & \text{otherwise}, \end{cases}
$$
where \(\mu\) is the Otsu threshold [2508.19300]. This mask marks signal-dominant regions and guides optimization toward genuine structural features rather than diffuse artifact regions [2508.19300].

Training is self-supervised and case-specific. The first loss term is a clean reconstruction loss over signal regions identified by the binary structure mask:
$$
\mathcal{L}_{\text{signal} = \frac{1}{N} \sum_{I}^N \left( I_{\text{clean} - \text{ReLU}(I_{\text{raw} - I_{\text{binary}) \right)^2,
$$
where only signal pixels contribute [2508.19300]. The second term is a 3D structural consistency loss based on total variation:
$$
\mathcal{L}_{TV}(X)=\sum_{i,j}\Bigl( |X_{i+1,j}-X_{i,j}| + |X_{i,j+1}-X_{i,j}| \Bigr).
$$
The final objective is
$$
\mathcal{L} = \mathcal{L}_{\text{signal} + \lambda \mathcal{L}_{TV},
$$
with \(\lambda=0.15\) in experiments [2508.19300]. The TV term is used to suppress isolated noise and encourage smooth spatial transitions, which the paper identifies as useful for maintaining 3D continuity [2508.19300].

## 5. Optimization protocol, architecture, and data resources

The optimization uses Adam with weight decay \(10^{-6}\), batch size 4096, up to 500,000 iterations, and a learning rate annealed linearly from \(2 \times 10^{-3}\) to \(2 \times 10^{-5}\) [2508.19300]. The MLP architecture matches NeRF’s design, with 8 fully connected layers, 256 channels per layer, and ReLU activations [2508.19300]. Training one \((256,356,160)\) volume took about 50,000 iterations and 20 minutes on a single RTX 4090 [2508.19300].

A notable contribution is the dataset construction. The paper describes this as the first paired 4D live-cell microscopy dataset for reconstruction evaluation, together with synthetic and unpaired public datasets for broader testing [2508.19300]. The paired real dataset comes from live *C. elegans* embryo imaging using a Leica SP5II confocal microscope with a 63×/1.4 NA objective [2508.19300]. Datasets 1, 2, and 3 are membrane-labeled with mCherry and excited at 561 nm, while dataset 4 is nucleus-labeled with GFP and excited at 488 nm [2508.19300]. The volumes have \(712 \times 512\) xy resolution and multiple z planes, later resampled to uniform \(0.18\,\mu m\) isotropic spacing and saved as 3D NIfTI volumes of size \((256,356,214)\) [2508.19300].

The dataset roles are differentiated as follows:

| Dataset | Role |
|---|---|
| Dataset 1 | Paired low- vs high-exposure real dataset |
| Dataset 2 | Synthetic/noise-augmented paired dataset |
| Dataset 3 | Longer time-resolved 4D series with prominent photo-induced artifacts |
| Dataset 4 | Synthetic/noise-augmented paired dataset |

Additional unpaired qualitative datasets include *C. elegans*, BAPE, zebrafish, and mouse/human embryo data [2508.19300]. For quantitative evaluation, the reported metrics are PSNR, SSIM, and LPIPS on datasets 1, 2, and 4, with mean and standard deviation over all samples [2508.19300].

## 6. Empirical performance, ablations, and comparative interpretation

The baselines listed in the paper are BM3D, Noise2Void, MM-BSN, and WBNS, together with ablations for baseline INR, w/o blind convolution, w/o structural amplification, and w/o TV loss [2508.19300]. Quantitatively, CellINR is reported to achieve the highest PSNR and SSIM and the lowest LPIPS across all paired datasets [2508.19300]. The paper further notes roughly a 30% improvement in mean values over mainstream methods on dataset 1, and about 10% improvement over WBNS [2508.19300].

WBNS is described as often the strongest competitor, but the paper argues that it can remove background artifacts at the cost of breaking structural continuity because wavelet decomposition can introduce serrated or discontinuous edges [2508.19300]. By contrast, CellINR is reported to better preserve continuity while eliminating pseudo-signals [2508.19300]. Qualitative evidence is described in terms of cleaner membranes and nuclei, smoother intensity profiles along selected line traces, and fewer jagged or blurred structures than competing methods [2508.19300].

The ablation results are presented as evidence that each design choice is functionally necessary. Baseline INR alone performs poorly and behaves nearly like an identity map [2508.19300]. Removing blind convolution hurts continuity and anti-aliasing [2508.19300]. Removing structural amplification weakens artifact suppression, especially for low-frequency background pseudo-signals [2508.19300]. Removing TV loss reduces suppression of high-frequency speckles and isolated noise [2508.19300]. Taken together, these observations support the paper’s claim that signal–artifact decoupling depends on combining spatial continuity, local neighborhood context, and frequency-aware coordinate encoding rather than relying on any one mechanism in isolation [2508.19300].

## 7. Scope, applications, and limitations

The main reported advantages of CellINR are that it is self-supervised, does not require pixel-aligned clean targets, explicitly models 3D continuity, leverages neighborhood context through blind convolution, and uses structural amplification to emphasize true biological regions [2508.19300]. It also supports arbitrary-resolution reconstruction through coordinate sampling, which is described as useful in microscopy where continuous spatial detail matters [2508.19300].

The paper suggests downstream value for 3D segmentation and morphological analysis during embryogenesis, stating that cleaner reconstructions can reduce misidentification and over-segmentation [2508.19300]. More broadly, the potential applications explicitly mentioned are long-term live-cell imaging, artifact suppression in embryogenesis studies, quantitative morphology analysis, cell segmentation, and biological imaging workflows in which preserving continuous structure is more important than merely smoothing noise [2508.19300].

The limitations are also explicit. The authors note that CellINR can misclassify regions when artifacts and true signals are ambiguous, so it is best suited to datasets where artifacts are clearly distinguishable from biological structures [2508.19300]. They also state that for multi-channel composite images, processing without separate single-channel data may mix fluorescence signals and degrade reconstruction quality [2508.19300]. Because the method is case-specific optimization, it may be slower than feed-forward models and may require per-volume training [2508.19300].

Within the broader methodological landscape, CellINR can be understood as a reconstruction framework that reframes artifact removal in 4D fluorescence microscopy as continuous implicit modeling rather than as ordinary denoising [2508.19300]. This suggests a shift in emphasis: from voxelwise filtering toward coordinate-based estimation constrained by local context and structural priors.

Source: https://www.emergentmind.com/topics/cellinr-framework