---
title: 'KonfAI: Configurable Deep Learning for Imaging'
url: https://www.emergentmind.com/topics/konfai
type: topic
---

# KonfAI: Configurable Deep Learning for Imaging

KonfAI is a modular, extensible, and fully configurable deep learning framework specifically designed for medical imaging tasks. It enables users to define complete training, inference, and evaluation workflows through structured YAML configuration files, without modifying the underlying code. Its declarative design is intended to enhance reproducibility, transparency, and experimental traceability while reducing development time, and it has been applied to segmentation, registration, and image synthesis tasks, including challenge settings. The framework is open source at `https://github.com/vboussot/KonfAI` [2508.09823].

## 1. System architecture

At its core, KonfAI is organized into five interchangeable modules, each responsible for one stage of the typical deep-learning pipeline. This modularization is not merely an implementation detail; it defines how experiment logic is partitioned and how configuration-time composition replaces code-level rewiring [2508.09823].

| Module | Responsibility |
|---|---|
| Data management layer | Data Loaders; Transforms & Augmentations |
| Model registry | Networks, Losses, Schedulers, Transforms, etc. |
| Training engine | Training loop orchestration |
| Inference engine | Reconstruction, TTA, ensembling, output formatting |
| Evaluation module | Metrics and JSON export |

The data management layer abstracts SimpleITK/HDF5 I/O and case-level grouping, and exposes configurable preprocessing, patch extraction, and data augmentation. The model registry provides a unified registry of Networks, Losses, Schedulers, Transforms, and related components, with each component referenced by its Python class path and constructor arguments. The training engine orchestrates the training loop, multi-GPU support, logging, checkpointing, learning-rate scheduling, and early stopping, and supports patch-based training in 2D, 2.5D, and 3D, as well as gradient accumulation and mixed precision. The inference engine handles patch-based reconstruction, test-time augmentation, model ensembling, and output formatting. The evaluation module compares predictions against ground truth, computes case-level and aggregate metrics, and exports JSON summaries.

A notable consequence of this structure is that the framework treats data handling, optimization, model instantiation, and evaluation as peer modules rather than hard-coded stages in a monolithic script. This suggests a design optimized for experiment composability rather than task-specific pipelines.

## 2. Declarative configuration and execution model

All modules communicate exclusively through three declarative YAML files—`Config.yml` for training, `Prediction.yml` for inference, and `Evaluation.yml` for metrics—loaded at runtime. KonfAI therefore adopts a fully YAML-driven interface in which experiment specification resides in configuration rather than in imperative training code [2508.09823].

The training workflow begins when the `TRAIN` command reads `Config.yml`, instantiates the Dataset, Model, Optimizer, Losses, and Scheduler, and saves training artifacts together with a snapshot of `Config.yml` into the workspace. The inference workflow begins when the `PREDICT` command reads `Prediction.yml`, reloads Model(s) and Dataset, executes patch-based or full-volume inference, optionally with test-time augmentation and ensembling, and writes outputs according to `outputs_dataset` definitions. The evaluation workflow begins when the `EVALUATE` command reads `Evaluation.yml`, loads ground truth and predictions, computes metrics, and emits per-case and global JSON files.

The corresponding command-line invocations are:

```bash
konfai train --config Config.yml
konfai predict --config Prediction.yml
konfai evaluate --config Evaluation.yml
```

This execution model makes the framework’s claim of reproducibility concrete. A workspace is not only a directory of checkpoints and predictions; it is also a serialized record of the experiment specification that produced them.

## 3. Built-in methodological strategies

KonfAI extends beyond standard pipelines by providing native abstractions for patch-based learning, test-time augmentation, model ensembling, and direct access to intermediate feature representations for deep supervision. It also supports complex multi-model training setups such as generative adversarial architectures [2508.09823].

Patch-based learning is configured in YAML through `patch_size`, `overlap`, and `pad_value`. The framework documentation gives the patch-sampling distribution as

$$
P(x) \;=\; \frac{1}{Z}\exp\Bigl(-\frac{\|x-\mu\|^2}{2\sigma^2}\Bigr).
$$

This makes the sampling policy an explicit part of the experiment description rather than an implicit preprocessing heuristic. Test-time augmentation is likewise declared in `Prediction.yml`, for example through horizontal flips and \(90^\circ\)/\(180^\circ\) rotations, with a specified reduction such as `Mean`. Model ensembling is configured by listing multiple checkpoints and weights, with the final output given as

$$
y = \sum_i w_i y_i.
$$

For deep supervision, KonfAI allows direct access to intermediate feature maps by naming submodules in `__init__()`. Multi-scale loss attachment is specified through `outputs_criterions`, and the overall loss is written as

$$
L \;=\; \sum_{s=0}^{S} \alpha_s \, L_s,
$$

where \(s\) indexes scale and \(\alpha_s\) are weights from the YAML schedule. In generative adversarial training, KonfAI can define both `Generator` and `Discriminator` inside the same configuration and combine their objectives through `GANLossCombiner`, with the summary expression \(L_{\mathrm{GAN}} = L_G + \lambda L_D\).

These mechanisms are significant because they are exposed as first-class configuration objects rather than bespoke extensions. A plausible implication is that KonfAI lowers the friction of comparing plain supervised setups against deep-supervision, ensembling, or adversarial variants within a common experiment substrate.

## 4. Registry design and extensibility

KonfAI’s extensibility is grounded in its registry-based architecture. Any custom Python file can be placed alongside the configuration files and referenced directly from YAML, and the same pattern applies to custom losses, transforms, schedulers, and models [2508.09823].

The documented example is a custom `FocalLoss` implemented in `CustomLoss.py` as a subclass of `konfai.criterions.Loss`, then referenced in YAML as `CustomLoss:FocalLoss` with constructor arguments such as `gamma`, `alpha`, `reduction`, and `is_loss: true`. Adding custom transforms and schedulers follows the same principle: create a subclass of `konfai.transforms.Transform` or `konfai.schedulers.Scheduler`, then reference `YourModule:YourClass` in YAML. The framework explicitly states that no core-code changes are required.

This design has two implications. First, it preserves the declarative contract even when introducing nonstandard components. Second, it permits local experimental specialization—custom losses, task-specific augmentations, or novel schedulers—without forking the framework’s internals. The practical guidance provided with the framework reinforces this pattern: develop custom modules in isolation, unit-test their forward methods before integrating, version control both YAML files and custom Python modules side-by-side, and use verbose logging or small synthetic datasets to validate the data path.

## 5. Reported task coverage and performance

KonfAI has been successfully applied to segmentation, registration, and image synthesis tasks, and the framework paper reports representative results across these categories [2508.09823].

For segmentation, reported use cases include brain tumor segmentation on BraTS with 3D patch training, achieving mean Dice \(\ge 0.90\) and Hausdorff distance \(< 5\) mm, and cardiac MRI with 2D U-Net, achieving mean Dice \(\approx 0.92\). For registration, the paper reports deformable registration on spine MRI with mean landmark error \(1.2\) mm. For image synthesis, it reports CT\(\rightarrow\)PET paired synthesis using cGAN with MAE \(\sim 0.15\) SUV units and SSIM \(> 0.85\).

The abstract further states that KonfAI has contributed to top-ranking results in several international medical imaging challenges. Within the framework’s own terms, these applications illustrate that the same declarative machinery can span dense prediction, geometric alignment, and generative synthesis without changing the core code path. That breadth is central to its identity: KonfAI is presented not as a single-task toolkit but as a configurable framework for heterogeneous medical-imaging workloads.

## 6. KonfAI in synthetic CT generation and registration-quality analysis

A detailed downstream use of KonfAI appears in the SynthRAD2025 challenge paper “Why Registration Quality Matters: Enhancing sCT Synthesis with IMPACT-Based Registration,” which implemented a unified MRI\(\rightarrow\)sCT and CBCT\(\rightarrow\)sCT pipeline in KonfAI [2510.21358].

In that study, the model was a 2.5D U-Net++ exactly as in Zhou et al. but with a ResNet-34 encoder from He et al. The KonfAI configuration registered `konfai.models.UnetPlusPlus` with a `torchvision.models.resnet34` encoder. Training used AdamW with initial learning rate \(0.001\) and weight decay \(10^{-2}\), combined with a StepLR scheduler halving the learning rate every \(25\,000\) steps. Batch size was \(32\), early stopping monitored validation MAE with patience of \(25\,000\) steps, and the study used nested five-fold cross-validation, followed by retraining five models on the full data and inference with test-time augmentation and ensembling. Final predictions leveraged test-time augmentation and five-fold ensembling, and the best model was selected based on validation MAE.

The data pipeline was equally explicit. Pre-alignment used either Elastix registration with mutual information or IMPACT registration. Body masking set voxels outside the provided body mask to \(-1\). Intensity normalization clipped CT to \([-1024,3071]\) and scaled to \([-1,1]\), clipped CBCT to \([\mathrm{min},99.5\text{th percentile}]\) in mask and scaled to \([-1,1]\), and standardized MRI inside mask to zero mean and unit variance. Patches were random \(320\times320\) for MRI or \(256\times256\) for CBCT with stride \(128\). The 2.5D input stacked \(\pm2\) slices for MRI, giving \(C_{\mathrm{in}}=5\), and \(\pm1\) slice for CBCT, giving \(C_{\mathrm{in}}=3\). The loss combined pixel-wise \(L1\) with IMPACT-Synth using \(\lambda=0.1\), and the trainer logged MAE, PSNR, and MS-SSIM.

The principal substantive finding concerned registration quality. On the local test sets, IMPACT-based registration achieved more accurate and anatomically consistent alignments than mutual-information-based registration, resulting in improved sCT synthesis with lower MAE and more realistic anatomical structures. On the public validation set, however, models trained with Elastix-aligned data achieved higher scores, reflecting a registration bias favoring alignment strategies consistent with the evaluation pipeline. The paper explicitly frames this as evidence that registration errors can propagate into supervised learning, influencing both training and evaluation, and potentially inflating performance metrics at the expense of anatomical fidelity. KonfAI’s role here was methodological: it provided a single, fully configurable pipeline in which registration strategy, loss function, and anatomical fine-tuning could be varied systematically.

## 7. Reproducibility, auditability, and naming ambiguity

Reproducibility in KonfAI is operationalized through a dedicated workspace per experiment that stores a config snapshot, Python module versions, git SHA, checkpoints, TensorBoard logs, predictions, and evaluation JSON, while automatic logging of hyperparameter changes is described as ensuring a full audit trail [2508.09823]. The framework also recommends `data_log` in YAML to visualize sample inputs, augmentations, and outputs in TensorBoard, and `--log_level DEBUG` to inspect instantiation steps. These features make experimental provenance a built-in artifact of execution rather than an afterthought.

A plausible source of confusion is that the label “KonfAI” also appears in a separate context in the CON-FOLD literature. The CON-FOLD paper concludes that its confidence-aware extension of FOLD-RM, with Wilson-interval rule confidences, confidence-based pruning, user-provided background knowledge, and the Inverse Brier Score, is “an ideal foundation for a ‘KonfAI’ system” [2408.07854]. This suggests that the term may be used aspirationally outside medical imaging. In the medical-imaging literature, however, KonfAI denotes the open-source, YAML-driven framework for training, inference, and evaluation workflows described above [2508.09823].

Source: https://www.emergentmind.com/topics/konfai