Papers
Topics
Authors
Recent
Search
2000 character limit reached

AutoRad-Lung: Radiomic AI for Nodule Malignancy

Updated 22 June 2026
  • AutoRad-Lung is a radiomic-guided vision-language model that fuses deep learning with 1,500 handcrafted CT radiomic features for lung nodule malignancy prediction.
  • Its multimodal pipeline integrates a frozen AIMv2 encoder and a CLIP-style text encoder, using a Meta-Net MLP to dynamically generate class-specific prompts.
  • The model achieves notable gains in accuracy and F1 scores, especially for ambiguous cases, highlighting its potential to enhance lung cancer screening.

AutoRad-Lung is a radiomic-guided prompting autoregressive vision-LLM designed for lung nodule malignancy prediction from computed tomography (CT) images (Khademi et al., 26 Mar 2025). It couples modern multimodal deep learning with handcrafted quantitative radiomic features to address core challenges in lung cancer screening, particularly the reliable classification of visually ambiguous pulmonary nodules.

1. Model Architecture and Multimodal Pipeline

AutoRad-Lung integrates a frozen, autoregressively pre-trained vision encoder (AIMv2, “large-patch14-224” variant) with a CLIP-style text encoder, conditionally prompted using radiomics. The input CT nodule slice (224×224224 \times 224) is decomposed into K=256K=256 non-overlapping 14×1414 \times 14 patches, each projected to a DD-dimensional embedding.

At pretraining, the joint model learns a causal autoregressive objective over image-patch and tokenized-text sequences: LAR=−∑t=1Klog⁡P(xt∣S<t)−∑u=1Tlog⁡P(wu∣S<K+u),\mathcal{L}_{\rm AR} = -\sum_{t=1}^K\log P(x_t \mid S_{<t}) - \sum_{u=1}^T\log P(w_u \mid S_{<K+u}), with S=[x1,...,xK,w1,...,wT]S = [x_1,...,x_K,w_1,...,w_T]. During downstream fine-tuning for malignancy prediction, the AIMv2 and CLIP encoders are frozen; only prompt-generation parameters are updated.

The prompt encoder receives “dynamic prompts” ti(r)=[v1(r),...,vM(r),ci]∈R(M+1)×Dt_i(\mathbf r) = [v_1(\mathbf r),...,v_M(\mathbf r),c_i] \in \mathbb{R}^{(M+1)\times D} for each class i∈{benign, malignant, unsure}i\in\{\text{benign, malignant, unsure}\}. Each prompt consists of MM context vectors adapted using radiomics (r∈R1500\mathbf r \in \mathbb{R}^{1500}) through a two-layer Meta-Net MLP, and a class token K=256K=2560. The class-specific text embedding K=256K=2561 is directly aligned with the AIMv2 image embedding K=256K=2562 via cosine similarity, yielding a softmax prediction: K=256K=2563 where K=256K=2564 and K=256K=2565 is a learned temperature parameter.

2. Radiomic Feature Extraction Methodology

AutoRad-Lung extracts 1,500 hand-crafted radiomic features from the central nodule slice using radiologist consensus masks (K=256K=2566 agreement). These features span several established families:

  • First-order statistics: mean, variance, skewness, entropy
  • Shape descriptors: volume K=256K=2567, surface area K=256K=2568, sphericity K=256K=2569
  • Texture metrics: GLCM (contrast, homogeneity, correlation), GLRLM (short/long-run emphasis), GLSZM (small-area emphasis), NGTDM (coarseness), GLDM (dependence non-uniformity)

These statistics are computed not just on the original image, but also on Laplacian-of-Gaussian, wavelet, square-root, logarithm, exponential, gradient, and LBP-filtered derivatives to form a comprehensive radiomic signature.

3. Conditional Context Optimization and Dynamic Prompting

For each test point, radiomic feature vector 14×1414 \times 140 is transformed by the Meta-Net MLP to obtain an offset 14×1414 \times 141, which is added to 14×1414 \times 142 learnable context tokens 14×1414 \times 143, producing prompt tokens 14×1414 \times 144. The resulting prompt 14×1414 \times 145 (for class 14×1414 \times 146) enables context-specific adaptation at inference, a methodological advance over CLIP-Lung and related VLMs, which limit prompt adaptation to training (Khademi et al., 26 Mar 2025).

The model aligns image and prompt embeddings in a joint space, with prediction performed as a softmax over cosine similarities, and the loss reduced to cross-entropy over classes: 14×1414 \times 147

4. Training Protocol and Inference Pipeline

Training is performed solely over prompt and Meta-Net parameters:

  • Frozen encoders: AIMv2 vision, CLIP-text (GPT-2–style, 12 blocks, 14×1414 \times 148)
  • Meta-Net: two-layer bottleneck MLP
  • Hyperparameters: SGD with momentum 14×1414 \times 149, weight decay DD0, batch size 64, initial learning rate DD1 (cosine-decay), epochs 30, DD2 context tokens
  • Hardware: Single NVIDIA RTX 3090 GPU

Inference steps:

  1. CT nodule preprocessing: central slice extraction, resampling to DD3, intensity normalization, cropping per consensus mask
  2. PyRadiomics computation of 1,500 features
  3. Generation of context-specific prompt via Meta-Net
  4. Image fed through frozen AIMv2 to produce DD4
  5. Text prompts for each class encoded by CLIP to yield DD5
  6. Prediction: class with maximal cosine similarity

5. Experimental Setup and Quantitative Results

Experiments use LIDC-IDRI annotated CTs (1,010 patients). Malignancy labels are trichotomized: benign (score DD62.5), malignant (DD73.5), unsure (DD8–DD9). Five-fold cross-validation yields the following mean performance:

Method Accuracy (%) Recall B/M/U (%) F1 B/M/U (%)
ResNet18 54.2 ± 0.6 72.2/64.4/29.0 62.0/61.3/36.6
UDM 54.6 ± 0.4 76.7/49.5/32.5 64.3/53.5/39.5
CLIP 56.6 ± 0.3 59.5/55.2/53.9 59.2/60.0/52.2
CoCoOp 56.8 ± 0.6 59.0/55.2/55.1 59.2/60.0/52.8
AIMv2 58.5 ± 0.3 62.5/43.6/51.3 55.2/45.6/52.3
CLIP-Lung 60.9 ± 0.4 67.5/60.9/53.4 64.4/66.3/54.1
AutoRad-Lung 64.6 ±1.7 75.3/65.6/62.3 49.5/60.6/71.6

AutoRad-Lung achieves a +3.7 percentage point (pp) gain in accuracy over CLIP-Lung, with the most pronounced gain (+17.5 pp) in F1-score for the “unsure” class (36.6%→71.6%). One-vs-rest ROC AUCs are all LAR=−∑t=1Klog⁡P(xt∣S<t)−∑u=1Tlog⁡P(wu∣S<K+u),\mathcal{L}_{\rm AR} = -\sum_{t=1}^K\log P(x_t \mid S_{<t}) - \sum_{u=1}^T\log P(w_u \mid S_{<K+u}),0. Peak performance occurs at LAR=−∑t=1Klog⁡P(xt∣S<t)−∑u=1Tlog⁡P(wu∣S<K+u),\mathcal{L}_{\rm AR} = -\sum_{t=1}^K\log P(x_t \mid S_{<t}) - \sum_{u=1}^T\log P(w_u \mid S_{<K+u}),1 context tokens; a greater LAR=−∑t=1Klog⁡P(xt∣S<t)−∑u=1Tlog⁡P(wu∣S<K+u),\mathcal{L}_{\rm AR} = -\sum_{t=1}^K\log P(x_t \mid S_{<t}) - \sum_{u=1}^T\log P(w_u \mid S_{<K+u}),2 results in overparameterization.

6. Context, Clinical Advantages, and Limitations

The methodology addresses three core limitations of prior vision-language approaches (Khademi et al., 26 Mar 2025):

  • Reduces dependence on subjective radiologist attribute annotations by using objective radiomics
  • Enables prompt-based textual guidance at inference via the conditional context optimization mechanism
  • Leverages prior knowledge within the vision encoder through transfer learning from large-scale multimodal autoregressive pretraining

Clinical significance is pronounced for visually ambiguous nodules, where radiomics+synthesis enhances sensitivity. A plausible implication is that integration of radiomics-derived context at inference may generalize to other low-contrast radiological tasks. However, radiomics extraction is currently limited to the central CT slice (not 3D), and 1,500-dimensional input increases compute burden. LIDC-IDRI is the only evaluation dataset; multi-center, multi-scanner validation is pending.

7. Future Directions

Suggested avenues include extending radiomic computation to volumetric (3D) features, learning sparse or attended radiomic subspaces, integrating patient-level clinical covariates (e.g., demographics, smoking status) into Meta-Net, and exploring end-to-end fine-tuning of both the vision and prompt modules contingent on larger, multi-institutional datasets.

AutoRad-Lung defines a new regime of radiology AI systems that align multimodal pre-trained representations with domain-specific quantitative biomarkers to improve clinically ambiguous prediction scenarios (Khademi et al., 26 Mar 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AutoRad-Lung.