Papers
Topics
Authors
Recent
Search
2000 character limit reached

Voxel-Level A3C Module for 3D Segmentation

Updated 14 January 2026
  • The paper presents a voxel-level A3C method that treats each voxel as an independent agent using composite rewards to correct noisy labels.
  • It leverages a shared encoder with multiple decoders to output segmentation, value, and policy predictions, ensuring stable asynchronous updates.
  • Empirical results demonstrate Dice score improvements of up to +3.72% on challenging datasets, confirming the module's robust performance.

The voxel-level Asynchronous Advantage Actor-Critic (vA3C) module is a reinforcement learning component formulated for robust 3D medical image segmentation in the presence of noisy annotations. It is a core innovation of the Staged Voxel-Level Deep Reinforcement Learning (SVL-DRL) framework, which frames each image voxel as an autonomous agent operating asynchronously in parallel. Unlike conventional sample-level or patch-level denoising approaches, vA3C exploits local policy adjustments driven by composite rewards that fuse segmentation accuracy metrics with anatomical constraints, thereby incrementally rectifying labeling inaccuracies at voxel resolution (Fu et al., 7 Jan 2026).

1. Voxel-wise Reinforcement Learning Agent Formulation

In the vA3C formulation, each voxel ii of a 3D image volume I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}—where N=H⋅W⋅DN = H \cdot W \cdot D—is modeled as an independent agent. The state of agent ii at step tt is denoted si(t)s_i^{(t)}, initialized to the voxel's raw intensity IiI_i and, in practice, set to the feature vector at ii within a global feature map F(t)F^{(t)} output by a shared encoder. This state captures both local and contextual neighborhood information. The collective system state is S(t)=(s1(t),…,sN(t))S^{(t)} = (s_1^{(t)}, \ldots, s_N^{(t)}), allowing all voxels to update policy-relevant representations simultaneously.

2. Discrete Action Space and Voxel Manipulations

Each voxel-agent samples from a finite action set I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}0:

  • I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}1: do nothing,
  • I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}2: enhance tissue/lesion,
  • I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}3: weaken tissue/lesion.

The selected action modulates the current voxel value via: I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}4 where I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}5. This formulation imposes controlled perturbations, simulating potential medical effects and facilitating error correction in the segmentation process.

3. Shared-Encoder and Multi-Decoder Architecture

The network backbone consists of a shared Swin-Unetr encoder, which generates hierarchical feature maps from the entire input volume I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}6. This is followed by three parallel decoder branches:

  • Segmentation head I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}7: produces probability map I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}8 for voxel-wise segmentation.
  • Value head I∈RH×W×DI \in \mathbb{R}^{H \times W \times D}9: estimates the expected return N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D0 for each voxel.
  • Policy head N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D1: outputs action logits for each voxel, parameterizing N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D2.

Encoder parameters are shared, while decoders for segmentation, value, and policy are independent, collectively parameterized by N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D3.

4. Voxel-Adaptive A3C: Workflow and Loss Formulations

The conventional A3C framework—where multiple actors interact with independent environments—is adapted such that every voxel is an agent in a shared environment, executed in parallel in each forward pass. Training proceeds in three main stages:

  1. Warmup: Pure supervised segmentation,
  2. Transition: Mixed Dice and value-MSE loss,
  3. Full RL: RL loss with segmentation, value, and policy terms.

At each RL step:

  • Actions are sampled per voxel: N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D4.
  • The action effect is applied to yield N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D5.
  • Step rewards N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D6 are computed (see Section 5).
  • N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D7-step returns are accumulated: N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D8, where N=Hâ‹…Wâ‹…DN = H \cdot W \cdot D9.
  • Advantages ii0 are used in policy gradients.
  • Gradients for segmentation, value, and policy loss are aggregated and used to update the parameters.

The staged loss functions are

ii1

Values for ii2 are set respectively to ii3.

5. Advantage Policy Gradient and Composite Reward Function

The advantage function follows the A3C standard: ii4. The policy gradient for voxel ii5 is

ii6

aggregated over all voxels and ii7-step windows.

The reward at each step is: ii8 where

ii9

and the anatomical constraint

tt0

with tt1 as the number of connected components in the binary mask, and tt2 the total variation of the segmentation. A lower tt3 enforces anatomical plausibility, penalizing fragmented or irregular outputs.

6. Empirical Validation: Ablation and Robustness under Noisy Annotations

Ablation studies (Tables 7–8 in the source) quantify vA3C’s contribution to segmentation accuracy under synthetic noise. For example, on the LA dataset with 50% SFDA-Noise, Dice scores improve from baseline tt4 to tt5 with full vA3C (+tt6); for Pancreas-CT with SFDA-Noise, performance increases from tt7 to tt8 (+tt9). Removing any of the RL training stages leads to further performance degradation. This empirically confirms that voxel-level asynchronous policy updates are the dominant factor for robust learning under annotation noise (Fu et al., 7 Jan 2026).

7. Theoretical Rationale for Voxel-Level Asynchronous Actor-Critic

Treating each voxel as an agent enables correction of mislabeled regions without requiring exclusion of entire volumes or patches. Asynchronous updates across spatially disjoint voxels decorrelate gradients, promoting stable optimization in high-dimensional settings. The composite reward, combining localized accuracy gain (Dice delta) and global anatomical regularity, incentivizes the model to incrementally ameliorate noisy labels. The staged training regimen ensures effective initialization, mitigating policy collapse. Overall, vA3C's granularity and parallel reinforcement maximize resilience to global annotation noise and accelerate convergence, outperforming conventional supervised baselines and sample-filtering methods (Fu et al., 7 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Voxel-Level Asynchronous Advantage Actor-Critic (vA3C) Module.