AutoRad-Lung: Radiomic AI for Nodule Malignancy
- AutoRad-Lung is a radiomic-guided vision-language model that fuses deep learning with 1,500 handcrafted CT radiomic features for lung nodule malignancy prediction.
- Its multimodal pipeline integrates a frozen AIMv2 encoder and a CLIP-style text encoder, using a Meta-Net MLP to dynamically generate class-specific prompts.
- The model achieves notable gains in accuracy and F1 scores, especially for ambiguous cases, highlighting its potential to enhance lung cancer screening.
AutoRad-Lung is a radiomic-guided prompting autoregressive vision-LLM designed for lung nodule malignancy prediction from computed tomography (CT) images (Khademi et al., 26 Mar 2025). It couples modern multimodal deep learning with handcrafted quantitative radiomic features to address core challenges in lung cancer screening, particularly the reliable classification of visually ambiguous pulmonary nodules.
1. Model Architecture and Multimodal Pipeline
AutoRad-Lung integrates a frozen, autoregressively pre-trained vision encoder (AIMv2, “large-patch14-224” variant) with a CLIP-style text encoder, conditionally prompted using radiomics. The input CT nodule slice () is decomposed into non-overlapping patches, each projected to a -dimensional embedding.
At pretraining, the joint model learns a causal autoregressive objective over image-patch and tokenized-text sequences: with . During downstream fine-tuning for malignancy prediction, the AIMv2 and CLIP encoders are frozen; only prompt-generation parameters are updated.
The prompt encoder receives “dynamic prompts” for each class . Each prompt consists of context vectors adapted using radiomics () through a two-layer Meta-Net MLP, and a class token 0. The class-specific text embedding 1 is directly aligned with the AIMv2 image embedding 2 via cosine similarity, yielding a softmax prediction: 3 where 4 and 5 is a learned temperature parameter.
2. Radiomic Feature Extraction Methodology
AutoRad-Lung extracts 1,500 hand-crafted radiomic features from the central nodule slice using radiologist consensus masks (6 agreement). These features span several established families:
- First-order statistics: mean, variance, skewness, entropy
- Shape descriptors: volume 7, surface area 8, sphericity 9
- Texture metrics: GLCM (contrast, homogeneity, correlation), GLRLM (short/long-run emphasis), GLSZM (small-area emphasis), NGTDM (coarseness), GLDM (dependence non-uniformity)
These statistics are computed not just on the original image, but also on Laplacian-of-Gaussian, wavelet, square-root, logarithm, exponential, gradient, and LBP-filtered derivatives to form a comprehensive radiomic signature.
3. Conditional Context Optimization and Dynamic Prompting
For each test point, radiomic feature vector 0 is transformed by the Meta-Net MLP to obtain an offset 1, which is added to 2 learnable context tokens 3, producing prompt tokens 4. The resulting prompt 5 (for class 6) enables context-specific adaptation at inference, a methodological advance over CLIP-Lung and related VLMs, which limit prompt adaptation to training (Khademi et al., 26 Mar 2025).
The model aligns image and prompt embeddings in a joint space, with prediction performed as a softmax over cosine similarities, and the loss reduced to cross-entropy over classes: 7
4. Training Protocol and Inference Pipeline
Training is performed solely over prompt and Meta-Net parameters:
- Frozen encoders: AIMv2 vision, CLIP-text (GPT-2–style, 12 blocks, 8)
- Meta-Net: two-layer bottleneck MLP
- Hyperparameters: SGD with momentum 9, weight decay 0, batch size 64, initial learning rate 1 (cosine-decay), epochs 30, 2 context tokens
- Hardware: Single NVIDIA RTX 3090 GPU
Inference steps:
- CT nodule preprocessing: central slice extraction, resampling to 3, intensity normalization, cropping per consensus mask
- PyRadiomics computation of 1,500 features
- Generation of context-specific prompt via Meta-Net
- Image fed through frozen AIMv2 to produce 4
- Text prompts for each class encoded by CLIP to yield 5
- Prediction: class with maximal cosine similarity
5. Experimental Setup and Quantitative Results
Experiments use LIDC-IDRI annotated CTs (1,010 patients). Malignancy labels are trichotomized: benign (score 62.5), malignant (73.5), unsure (8–9). Five-fold cross-validation yields the following mean performance:
| Method | Accuracy (%) | Recall B/M/U (%) | F1 B/M/U (%) |
|---|---|---|---|
| ResNet18 | 54.2 ± 0.6 | 72.2/64.4/29.0 | 62.0/61.3/36.6 |
| UDM | 54.6 ± 0.4 | 76.7/49.5/32.5 | 64.3/53.5/39.5 |
| CLIP | 56.6 ± 0.3 | 59.5/55.2/53.9 | 59.2/60.0/52.2 |
| CoCoOp | 56.8 ± 0.6 | 59.0/55.2/55.1 | 59.2/60.0/52.8 |
| AIMv2 | 58.5 ± 0.3 | 62.5/43.6/51.3 | 55.2/45.6/52.3 |
| CLIP-Lung | 60.9 ± 0.4 | 67.5/60.9/53.4 | 64.4/66.3/54.1 |
| AutoRad-Lung | 64.6 ±1.7 | 75.3/65.6/62.3 | 49.5/60.6/71.6 |
AutoRad-Lung achieves a +3.7 percentage point (pp) gain in accuracy over CLIP-Lung, with the most pronounced gain (+17.5 pp) in F1-score for the “unsure” class (36.6%→71.6%). One-vs-rest ROC AUCs are all 0. Peak performance occurs at 1 context tokens; a greater 2 results in overparameterization.
6. Context, Clinical Advantages, and Limitations
The methodology addresses three core limitations of prior vision-language approaches (Khademi et al., 26 Mar 2025):
- Reduces dependence on subjective radiologist attribute annotations by using objective radiomics
- Enables prompt-based textual guidance at inference via the conditional context optimization mechanism
- Leverages prior knowledge within the vision encoder through transfer learning from large-scale multimodal autoregressive pretraining
Clinical significance is pronounced for visually ambiguous nodules, where radiomics+synthesis enhances sensitivity. A plausible implication is that integration of radiomics-derived context at inference may generalize to other low-contrast radiological tasks. However, radiomics extraction is currently limited to the central CT slice (not 3D), and 1,500-dimensional input increases compute burden. LIDC-IDRI is the only evaluation dataset; multi-center, multi-scanner validation is pending.
7. Future Directions
Suggested avenues include extending radiomic computation to volumetric (3D) features, learning sparse or attended radiomic subspaces, integrating patient-level clinical covariates (e.g., demographics, smoking status) into Meta-Net, and exploring end-to-end fine-tuning of both the vision and prompt modules contingent on larger, multi-institutional datasets.
AutoRad-Lung defines a new regime of radiology AI systems that align multimodal pre-trained representations with domain-specific quantitative biomarkers to improve clinically ambiguous prediction scenarios (Khademi et al., 26 Mar 2025).