Papers
Topics
Authors
Recent
Search
2000 character limit reached

KonfAI: Configurable Deep Learning for Imaging

Updated 8 July 2026
  • KonfAI is a modular and configurable deep learning framework for medical imaging tasks, enabling YAML-driven definitions of training, inference, and evaluation.
  • Its architecture partitions the workflow into interchangeable modules—data management, model registry, training engine, inference engine, and evaluation—enhancing reproducibility and composability.
  • The framework supports advanced techniques such as patch-based learning, test-time augmentation, and model ensembling, while allowing seamless integration of custom modules.

KonfAI is a modular, extensible, and fully configurable deep learning framework specifically designed for medical imaging tasks. It enables users to define complete training, inference, and evaluation workflows through structured YAML configuration files, without modifying the underlying code. Its declarative design is intended to enhance reproducibility, transparency, and experimental traceability while reducing development time, and it has been applied to segmentation, registration, and image synthesis tasks, including challenge settings. The framework is open source at https://github.com/vboussot/KonfAI (Boussot et al., 13 Aug 2025).

1. System architecture

At its core, KonfAI is organized into five interchangeable modules, each responsible for one stage of the typical deep-learning pipeline. This modularization is not merely an implementation detail; it defines how experiment logic is partitioned and how configuration-time composition replaces code-level rewiring (Boussot et al., 13 Aug 2025).

Module Responsibility
Data management layer Data Loaders; Transforms & Augmentations
Model registry Networks, Losses, Schedulers, Transforms, etc.
Training engine Training loop orchestration
Inference engine Reconstruction, TTA, ensembling, output formatting
Evaluation module Metrics and JSON export

The data management layer abstracts SimpleITK/HDF5 I/O and case-level grouping, and exposes configurable preprocessing, patch extraction, and data augmentation. The model registry provides a unified registry of Networks, Losses, Schedulers, Transforms, and related components, with each component referenced by its Python class path and constructor arguments. The training engine orchestrates the training loop, multi-GPU support, logging, checkpointing, learning-rate scheduling, and early stopping, and supports patch-based training in 2D, 2.5D, and 3D, as well as gradient accumulation and mixed precision. The inference engine handles patch-based reconstruction, test-time augmentation, model ensembling, and output formatting. The evaluation module compares predictions against ground truth, computes case-level and aggregate metrics, and exports JSON summaries.

A notable consequence of this structure is that the framework treats data handling, optimization, model instantiation, and evaluation as peer modules rather than hard-coded stages in a monolithic script. This suggests a design optimized for experiment composability rather than task-specific pipelines.

2. Declarative configuration and execution model

All modules communicate exclusively through three declarative YAML files—Config.yml for training, Prediction.yml for inference, and Evaluation.yml for metrics—loaded at runtime. KonfAI therefore adopts a fully YAML-driven interface in which experiment specification resides in configuration rather than in imperative training code (Boussot et al., 13 Aug 2025).

The training workflow begins when the TRAIN command reads Config.yml, instantiates the Dataset, Model, Optimizer, Losses, and Scheduler, and saves training artifacts together with a snapshot of Config.yml into the workspace. The inference workflow begins when the PREDICT command reads Prediction.yml, reloads Model(s) and Dataset, executes patch-based or full-volume inference, optionally with test-time augmentation and ensembling, and writes outputs according to outputs_dataset definitions. The evaluation workflow begins when the EVALUATE command reads Evaluation.yml, loads ground truth and predictions, computes metrics, and emits per-case and global JSON files.

The corresponding command-line invocations are:

y=∑iwiyi.y = \sum_i w_i y_i.6

This execution model makes the framework’s claim of reproducibility concrete. A workspace is not only a directory of checkpoints and predictions; it is also a serialized record of the experiment specification that produced them.

3. Built-in methodological strategies

KonfAI extends beyond standard pipelines by providing native abstractions for patch-based learning, test-time augmentation, model ensembling, and direct access to intermediate feature representations for deep supervision. It also supports complex multi-model training setups such as generative adversarial architectures (Boussot et al., 13 Aug 2025).

Patch-based learning is configured in YAML through patch_size, overlap, and pad_value. The framework documentation gives the patch-sampling distribution as

P(x)  =  1Zexp⁡(−∥x−μ∥22σ2).P(x) \;=\; \frac{1}{Z}\exp\Bigl(-\frac{\|x-\mu\|^2}{2\sigma^2}\Bigr).

This makes the sampling policy an explicit part of the experiment description rather than an implicit preprocessing heuristic. Test-time augmentation is likewise declared in Prediction.yml, for example through horizontal flips and 90∘90^\circ/180∘180^\circ rotations, with a specified reduction such as Mean. Model ensembling is configured by listing multiple checkpoints and weights, with the final output given as

y=∑iwiyi.y = \sum_i w_i y_i.

For deep supervision, KonfAI allows direct access to intermediate feature maps by naming submodules in __init__(). Multi-scale loss attachment is specified through outputs_criterions, and the overall loss is written as

L  =  ∑s=0Sαs Ls,L \;=\; \sum_{s=0}^{S} \alpha_s \, L_s,

where ss indexes scale and αs\alpha_s are weights from the YAML schedule. In generative adversarial training, KonfAI can define both Generator and Discriminator inside the same configuration and combine their objectives through GANLossCombiner, with the summary expression LGAN=LG+λLDL_{\mathrm{GAN}} = L_G + \lambda L_D.

These mechanisms are significant because they are exposed as first-class configuration objects rather than bespoke extensions. A plausible implication is that KonfAI lowers the friction of comparing plain supervised setups against deep-supervision, ensembling, or adversarial variants within a common experiment substrate.

4. Registry design and extensibility

KonfAI’s extensibility is grounded in its registry-based architecture. Any custom Python file can be placed alongside the configuration files and referenced directly from YAML, and the same pattern applies to custom losses, transforms, schedulers, and models (Boussot et al., 13 Aug 2025).

The documented example is a custom FocalLoss implemented in CustomLoss.py as a subclass of konfai.criterions.Loss, then referenced in YAML as CustomLoss:FocalLoss with constructor arguments such as gamma, alpha, reduction, and is_loss: true. Adding custom transforms and schedulers follows the same principle: create a subclass of konfai.transforms.Transform or konfai.schedulers.Scheduler, then reference YourModule:YourClass in YAML. The framework explicitly states that no core-code changes are required.

This design has two implications. First, it preserves the declarative contract even when introducing nonstandard components. Second, it permits local experimental specialization—custom losses, task-specific augmentations, or novel schedulers—without forking the framework’s internals. The practical guidance provided with the framework reinforces this pattern: develop custom modules in isolation, unit-test their forward methods before integrating, version control both YAML files and custom Python modules side-by-side, and use verbose logging or small synthetic datasets to validate the data path.

5. Reported task coverage and performance

KonfAI has been successfully applied to segmentation, registration, and image synthesis tasks, and the framework paper reports representative results across these categories (Boussot et al., 13 Aug 2025).

For segmentation, reported use cases include brain tumor segmentation on BraTS with 3D patch training, achieving mean Dice ≥0.90\ge 0.90 and Hausdorff distance <5< 5 mm, and cardiac MRI with 2D U-Net, achieving mean Dice 90∘90^\circ0. For registration, the paper reports deformable registration on spine MRI with mean landmark error 90∘90^\circ1 mm. For image synthesis, it reports CT90∘90^\circ2PET paired synthesis using cGAN with MAE 90∘90^\circ3 SUV units and SSIM 90∘90^\circ4.

The abstract further states that KonfAI has contributed to top-ranking results in several international medical imaging challenges. Within the framework’s own terms, these applications illustrate that the same declarative machinery can span dense prediction, geometric alignment, and generative synthesis without changing the core code path. That breadth is central to its identity: KonfAI is presented not as a single-task toolkit but as a configurable framework for heterogeneous medical-imaging workloads.

6. KonfAI in synthetic CT generation and registration-quality analysis

A detailed downstream use of KonfAI appears in the SynthRAD2025 challenge paper “Why Registration Quality Matters: Enhancing sCT Synthesis with IMPACT-Based Registration,” which implemented a unified MRI90∘90^\circ5sCT and CBCT90∘90^\circ6sCT pipeline in KonfAI (Boussot et al., 24 Oct 2025).

In that study, the model was a 2.5D U-Net++ exactly as in Zhou et al. but with a ResNet-34 encoder from He et al. The KonfAI configuration registered konfai.models.UnetPlusPlus with a torchvision.models.resnet34 encoder. Training used AdamW with initial learning rate 90∘90^\circ7 and weight decay 90∘90^\circ8, combined with a StepLR scheduler halving the learning rate every 90∘90^\circ9 steps. Batch size was 180∘180^\circ0, early stopping monitored validation MAE with patience of 180∘180^\circ1 steps, and the study used nested five-fold cross-validation, followed by retraining five models on the full data and inference with test-time augmentation and ensembling. Final predictions leveraged test-time augmentation and five-fold ensembling, and the best model was selected based on validation MAE.

The data pipeline was equally explicit. Pre-alignment used either Elastix registration with mutual information or IMPACT registration. Body masking set voxels outside the provided body mask to 180∘180^\circ2. Intensity normalization clipped CT to 180∘180^\circ3 and scaled to 180∘180^\circ4, clipped CBCT to 180∘180^\circ5 in mask and scaled to 180∘180^\circ6, and standardized MRI inside mask to zero mean and unit variance. Patches were random 180∘180^\circ7 for MRI or 180∘180^\circ8 for CBCT with stride 180∘180^\circ9. The 2.5D input stacked y=∑iwiyi.y = \sum_i w_i y_i.0 slices for MRI, giving y=∑iwiyi.y = \sum_i w_i y_i.1, and y=∑iwiyi.y = \sum_i w_i y_i.2 slice for CBCT, giving y=∑iwiyi.y = \sum_i w_i y_i.3. The loss combined pixel-wise y=∑iwiyi.y = \sum_i w_i y_i.4 with IMPACT-Synth using y=∑iwiyi.y = \sum_i w_i y_i.5, and the trainer logged MAE, PSNR, and MS-SSIM.

The principal substantive finding concerned registration quality. On the local test sets, IMPACT-based registration achieved more accurate and anatomically consistent alignments than mutual-information-based registration, resulting in improved sCT synthesis with lower MAE and more realistic anatomical structures. On the public validation set, however, models trained with Elastix-aligned data achieved higher scores, reflecting a registration bias favoring alignment strategies consistent with the evaluation pipeline. The paper explicitly frames this as evidence that registration errors can propagate into supervised learning, influencing both training and evaluation, and potentially inflating performance metrics at the expense of anatomical fidelity. KonfAI’s role here was methodological: it provided a single, fully configurable pipeline in which registration strategy, loss function, and anatomical fine-tuning could be varied systematically.

7. Reproducibility, auditability, and naming ambiguity

Reproducibility in KonfAI is operationalized through a dedicated workspace per experiment that stores a config snapshot, Python module versions, git SHA, checkpoints, TensorBoard logs, predictions, and evaluation JSON, while automatic logging of hyperparameter changes is described as ensuring a full audit trail (Boussot et al., 13 Aug 2025). The framework also recommends data_log in YAML to visualize sample inputs, augmentations, and outputs in TensorBoard, and --log_level DEBUG to inspect instantiation steps. These features make experimental provenance a built-in artifact of execution rather than an afterthought.

A plausible source of confusion is that the label “KonfAI” also appears in a separate context in the CON-FOLD literature. The CON-FOLD paper concludes that its confidence-aware extension of FOLD-RM, with Wilson-interval rule confidences, confidence-based pruning, user-provided background knowledge, and the Inverse Brier Score, is “an ideal foundation for a ‘KonfAI’ system” (McGinness et al., 2024). This suggests that the term may be used aspirationally outside medical imaging. In the medical-imaging literature, however, KonfAI denotes the open-source, YAML-driven framework for training, inference, and evaluation workflows described above (Boussot et al., 13 Aug 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to KonfAI.