KonfAI: Configurable Deep Learning for Imaging
- KonfAI is a modular and configurable deep learning framework for medical imaging tasks, enabling YAML-driven definitions of training, inference, and evaluation.
- Its architecture partitions the workflow into interchangeable modules—data management, model registry, training engine, inference engine, and evaluation—enhancing reproducibility and composability.
- The framework supports advanced techniques such as patch-based learning, test-time augmentation, and model ensembling, while allowing seamless integration of custom modules.
KonfAI is a modular, extensible, and fully configurable deep learning framework specifically designed for medical imaging tasks. It enables users to define complete training, inference, and evaluation workflows through structured YAML configuration files, without modifying the underlying code. Its declarative design is intended to enhance reproducibility, transparency, and experimental traceability while reducing development time, and it has been applied to segmentation, registration, and image synthesis tasks, including challenge settings. The framework is open source at https://github.com/vboussot/KonfAI (Boussot et al., 13 Aug 2025).
1. System architecture
At its core, KonfAI is organized into five interchangeable modules, each responsible for one stage of the typical deep-learning pipeline. This modularization is not merely an implementation detail; it defines how experiment logic is partitioned and how configuration-time composition replaces code-level rewiring (Boussot et al., 13 Aug 2025).
| Module | Responsibility |
|---|---|
| Data management layer | Data Loaders; Transforms & Augmentations |
| Model registry | Networks, Losses, Schedulers, Transforms, etc. |
| Training engine | Training loop orchestration |
| Inference engine | Reconstruction, TTA, ensembling, output formatting |
| Evaluation module | Metrics and JSON export |
The data management layer abstracts SimpleITK/HDF5 I/O and case-level grouping, and exposes configurable preprocessing, patch extraction, and data augmentation. The model registry provides a unified registry of Networks, Losses, Schedulers, Transforms, and related components, with each component referenced by its Python class path and constructor arguments. The training engine orchestrates the training loop, multi-GPU support, logging, checkpointing, learning-rate scheduling, and early stopping, and supports patch-based training in 2D, 2.5D, and 3D, as well as gradient accumulation and mixed precision. The inference engine handles patch-based reconstruction, test-time augmentation, model ensembling, and output formatting. The evaluation module compares predictions against ground truth, computes case-level and aggregate metrics, and exports JSON summaries.
A notable consequence of this structure is that the framework treats data handling, optimization, model instantiation, and evaluation as peer modules rather than hard-coded stages in a monolithic script. This suggests a design optimized for experiment composability rather than task-specific pipelines.
2. Declarative configuration and execution model
All modules communicate exclusively through three declarative YAML files—Config.yml for training, Prediction.yml for inference, and Evaluation.yml for metrics—loaded at runtime. KonfAI therefore adopts a fully YAML-driven interface in which experiment specification resides in configuration rather than in imperative training code (Boussot et al., 13 Aug 2025).
The training workflow begins when the TRAIN command reads Config.yml, instantiates the Dataset, Model, Optimizer, Losses, and Scheduler, and saves training artifacts together with a snapshot of Config.yml into the workspace. The inference workflow begins when the PREDICT command reads Prediction.yml, reloads Model(s) and Dataset, executes patch-based or full-volume inference, optionally with test-time augmentation and ensembling, and writes outputs according to outputs_dataset definitions. The evaluation workflow begins when the EVALUATE command reads Evaluation.yml, loads ground truth and predictions, computes metrics, and emits per-case and global JSON files.
The corresponding command-line invocations are:
6
This execution model makes the framework’s claim of reproducibility concrete. A workspace is not only a directory of checkpoints and predictions; it is also a serialized record of the experiment specification that produced them.
3. Built-in methodological strategies
KonfAI extends beyond standard pipelines by providing native abstractions for patch-based learning, test-time augmentation, model ensembling, and direct access to intermediate feature representations for deep supervision. It also supports complex multi-model training setups such as generative adversarial architectures (Boussot et al., 13 Aug 2025).
Patch-based learning is configured in YAML through patch_size, overlap, and pad_value. The framework documentation gives the patch-sampling distribution as
This makes the sampling policy an explicit part of the experiment description rather than an implicit preprocessing heuristic. Test-time augmentation is likewise declared in Prediction.yml, for example through horizontal flips and / rotations, with a specified reduction such as Mean. Model ensembling is configured by listing multiple checkpoints and weights, with the final output given as
For deep supervision, KonfAI allows direct access to intermediate feature maps by naming submodules in __init__(). Multi-scale loss attachment is specified through outputs_criterions, and the overall loss is written as
where indexes scale and are weights from the YAML schedule. In generative adversarial training, KonfAI can define both Generator and Discriminator inside the same configuration and combine their objectives through GANLossCombiner, with the summary expression .
These mechanisms are significant because they are exposed as first-class configuration objects rather than bespoke extensions. A plausible implication is that KonfAI lowers the friction of comparing plain supervised setups against deep-supervision, ensembling, or adversarial variants within a common experiment substrate.
4. Registry design and extensibility
KonfAI’s extensibility is grounded in its registry-based architecture. Any custom Python file can be placed alongside the configuration files and referenced directly from YAML, and the same pattern applies to custom losses, transforms, schedulers, and models (Boussot et al., 13 Aug 2025).
The documented example is a custom FocalLoss implemented in CustomLoss.py as a subclass of konfai.criterions.Loss, then referenced in YAML as CustomLoss:FocalLoss with constructor arguments such as gamma, alpha, reduction, and is_loss: true. Adding custom transforms and schedulers follows the same principle: create a subclass of konfai.transforms.Transform or konfai.schedulers.Scheduler, then reference YourModule:YourClass in YAML. The framework explicitly states that no core-code changes are required.
This design has two implications. First, it preserves the declarative contract even when introducing nonstandard components. Second, it permits local experimental specialization—custom losses, task-specific augmentations, or novel schedulers—without forking the framework’s internals. The practical guidance provided with the framework reinforces this pattern: develop custom modules in isolation, unit-test their forward methods before integrating, version control both YAML files and custom Python modules side-by-side, and use verbose logging or small synthetic datasets to validate the data path.
5. Reported task coverage and performance
KonfAI has been successfully applied to segmentation, registration, and image synthesis tasks, and the framework paper reports representative results across these categories (Boussot et al., 13 Aug 2025).
For segmentation, reported use cases include brain tumor segmentation on BraTS with 3D patch training, achieving mean Dice and Hausdorff distance mm, and cardiac MRI with 2D U-Net, achieving mean Dice 0. For registration, the paper reports deformable registration on spine MRI with mean landmark error 1 mm. For image synthesis, it reports CT2PET paired synthesis using cGAN with MAE 3 SUV units and SSIM 4.
The abstract further states that KonfAI has contributed to top-ranking results in several international medical imaging challenges. Within the framework’s own terms, these applications illustrate that the same declarative machinery can span dense prediction, geometric alignment, and generative synthesis without changing the core code path. That breadth is central to its identity: KonfAI is presented not as a single-task toolkit but as a configurable framework for heterogeneous medical-imaging workloads.
6. KonfAI in synthetic CT generation and registration-quality analysis
A detailed downstream use of KonfAI appears in the SynthRAD2025 challenge paper “Why Registration Quality Matters: Enhancing sCT Synthesis with IMPACT-Based Registration,” which implemented a unified MRI5sCT and CBCT6sCT pipeline in KonfAI (Boussot et al., 24 Oct 2025).
In that study, the model was a 2.5D U-Net++ exactly as in Zhou et al. but with a ResNet-34 encoder from He et al. The KonfAI configuration registered konfai.models.UnetPlusPlus with a torchvision.models.resnet34 encoder. Training used AdamW with initial learning rate 7 and weight decay 8, combined with a StepLR scheduler halving the learning rate every 9 steps. Batch size was 0, early stopping monitored validation MAE with patience of 1 steps, and the study used nested five-fold cross-validation, followed by retraining five models on the full data and inference with test-time augmentation and ensembling. Final predictions leveraged test-time augmentation and five-fold ensembling, and the best model was selected based on validation MAE.
The data pipeline was equally explicit. Pre-alignment used either Elastix registration with mutual information or IMPACT registration. Body masking set voxels outside the provided body mask to 2. Intensity normalization clipped CT to 3 and scaled to 4, clipped CBCT to 5 in mask and scaled to 6, and standardized MRI inside mask to zero mean and unit variance. Patches were random 7 for MRI or 8 for CBCT with stride 9. The 2.5D input stacked 0 slices for MRI, giving 1, and 2 slice for CBCT, giving 3. The loss combined pixel-wise 4 with IMPACT-Synth using 5, and the trainer logged MAE, PSNR, and MS-SSIM.
The principal substantive finding concerned registration quality. On the local test sets, IMPACT-based registration achieved more accurate and anatomically consistent alignments than mutual-information-based registration, resulting in improved sCT synthesis with lower MAE and more realistic anatomical structures. On the public validation set, however, models trained with Elastix-aligned data achieved higher scores, reflecting a registration bias favoring alignment strategies consistent with the evaluation pipeline. The paper explicitly frames this as evidence that registration errors can propagate into supervised learning, influencing both training and evaluation, and potentially inflating performance metrics at the expense of anatomical fidelity. KonfAI’s role here was methodological: it provided a single, fully configurable pipeline in which registration strategy, loss function, and anatomical fine-tuning could be varied systematically.
7. Reproducibility, auditability, and naming ambiguity
Reproducibility in KonfAI is operationalized through a dedicated workspace per experiment that stores a config snapshot, Python module versions, git SHA, checkpoints, TensorBoard logs, predictions, and evaluation JSON, while automatic logging of hyperparameter changes is described as ensuring a full audit trail (Boussot et al., 13 Aug 2025). The framework also recommends data_log in YAML to visualize sample inputs, augmentations, and outputs in TensorBoard, and --log_level DEBUG to inspect instantiation steps. These features make experimental provenance a built-in artifact of execution rather than an afterthought.
A plausible source of confusion is that the label “KonfAI” also appears in a separate context in the CON-FOLD literature. The CON-FOLD paper concludes that its confidence-aware extension of FOLD-RM, with Wilson-interval rule confidences, confidence-based pruning, user-provided background knowledge, and the Inverse Brier Score, is “an ideal foundation for a ‘KonfAI’ system” (McGinness et al., 2024). This suggests that the term may be used aspirationally outside medical imaging. In the medical-imaging literature, however, KonfAI denotes the open-source, YAML-driven framework for training, inference, and evaluation workflows described above (Boussot et al., 13 Aug 2025).