Papers
Topics
Authors
Recent
Search
2000 character limit reached

Nested MIL with Attention

Updated 20 January 2026
  • The paper introduces a nested MIL framework that employs dedicated attention mechanisms at every hierarchy level to enhance prediction accuracy and interpretability.
  • NMIA is a hierarchical model that uses multi-level feature embedding and aggregation to capture complex dependencies in weakly supervised, bag-of-bags data.
  • Empirical evaluations demonstrate that NMIA outperforms traditional MIL approaches, particularly in tasks requiring structured latent label inference and rule-based aggregation.

Nested Multiple Instance Learning with Attention (NMIA) extends the canonical Multiple Instance Learning (MIL) paradigm to address weakly supervised problems with complex hierarchical structures, where only bag-of-bags labels are available and neither instance nor inner-bag labels are observed. NMIA introduces JJ levels of bag nesting and employs dedicated attention mechanisms at each level. This framework enables not only accurate prediction of the outermost bag labels but also interpretable soft predictions of latent labels at lower levels. The original model formulation and empirical analysis are detailed in "Nested Multiple Instance Learning with Attention Mechanisms" (Fuster et al., 2021).

1. Hierarchical Weak Supervision and Formal Setup

NMIA formalizes a setting where only the label y∈{0,1}y \in \{0,1\} of a single outermost bag XX is observed, but the data structure is intrinsically hierarchical:

  • Level 1 (Innermost): Instances x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D, grouped into inner-bags.
  • Levels 2...J–1: Each level jj comprises bags Xj,kX_{j,k} of elements from level j−1j-1.
  • Level J (Outermost): The top-level bag XJ,1X_{J,1} contains inner-bags XJ−1,kX_{J-1,k}, for k=1...KJ−1k=1...K_{J-1}.

Notation:

  • y∈{0,1}y \in \{0,1\}0: y∈{0,1}y \in \{0,1\}1-th element of y∈{0,1}y \in \{0,1\}2-th bag at level y∈{0,1}y \in \{0,1\}3 (y∈{0,1}y \in \{0,1\}4 for instance; y∈{0,1}y \in \{0,1\}5 for embedding of sub-bag if y∈{0,1}y \in \{0,1\}6).
  • y∈{0,1}y \in \{0,1\}7: y∈{0,1}y \in \{0,1\}8-th bag at level y∈{0,1}y \in \{0,1\}9, XX0 elements.
  • XX1.
  • XX2: latent label of XX3 (not observed).
  • Under standard MIL (XX4), XX5.

This nested organization generalizes MIL such that models can capture complex dependencies, like grouping similar instances or enforcing relational bag rules.

2. Model Architecture and Attention Mechanisms

NMIA employs a multi-tiered process for representation and aggregation, parameterized as follows:

2.1 Instance-level Feature Embedding

Each raw instance XX6 is embedded:

XX7

where XX8 is typically a CNN or MLP.

2.2 Attention from Instance to Inner-bag

Attention scores for each instance in its inner-bag:

XX9

x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D0

with x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D1, x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D2.

A gated-attention variant is also considered:

x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D3

with x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D4, x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D5 element-wise, x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D6 sigmoid.

2.3 Inner-Bag Representation Aggregation

Weighted sum for each inner-bag:

x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D7

2.4 Attention from Inner-bag to Outer-bag

Aggregation to the outer-bag:

x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D8

x1,k,l∈RDx_{1,k,l} \in \mathbb{R}^D9

with jj0, jj1.

Final bag-of-bags embedding:

jj2

2.5 Classification Head

Prediction via:

jj3

where jj4 is the predicted probability.

3. Training Objective and Optimization

The model is trained end-to-end with combined parameters jj5, minimizing binary cross-entropy on outer-bag labels and optional jj6 regularization:

jj7

Early stopping is typically applied using a held-out validation set.

The full process from instance embedding to nested attention aggregation is differentiable, amenable to optimization by SGD or Adam.

4. Latent Label Prediction via Hierarchical Attention

Although supervision is available only at the outer bag level, NMIA leverages nested attention for latent label inference:

  • Instance-level score jj8: Measures likelihood that instance jj9 is positive within its inner-bag; thresholding Xj,kX_{j,k}0 enables latent positive assignment Xj,kX_{j,k}1.
  • Inner-bag score Xj,kX_{j,k}2: Indicates inner-bag Xj,kX_{j,k}3 contribution to the positive outer-label; thresholding Xj,kX_{j,k}4 yields latent positive inner-bag assignment Xj,kX_{j,k}5.

This nested inference enables partial recovery of latent structure, as shown in medical whole-slide imaging (WSI) examples where attention highlights candidate lesions and regions.

5. Computational Workflow

The NMIA training/inference procedure is directly expressed in the following pseudocode:

j−1j-10

At inference, the same forward pass provides Xj,kX_{j,k}6 and attention maps Xj,kX_{j,k}7, supporting both outer prediction and interpretability for inner structure via attention thresholding.

6. Empirical Evaluation and Comparative Results

NMIA was evaluated on two-level (MNIST, PCAM) and three-level (MNIST "odd-only" rule) benchmarks, compared with alternative MIL architectures:

Dataset/Experiment MI MIA NMI NMIA
MNIST Exp1 (single-instance→bag) 0.929 0.957 0.923 0.959
MNIST Exp2 (≥2 positives in same inner-bag) 0.345 0.472 0.855 0.921
MNIST Exp3 (3-level "odd-only" rule) N/A N/A 0.556 0.836
PCAM Exp1 (standard MIL) 0.957 0.973 0.964 0.978
PCAM Exp2 (≥2 metastatic patches/region) 0.290 0.286 0.700 0.734
  • In easy tasks (Exp1), all models perform well, with NMIA slightly outperforming alternatives.
  • For rule-based tasks requiring the grouping of positives (Exp2), conventional MI/MIA architectures fail, while NMI and NMIA model the required relations, with NMIA achieving superior F1.
  • The three-level hierarchy (Exp3) demonstrates only the NMIA architecture's capacity to learn complex hierarchical rules (e.g., aggregating presence/absence across nested levels).
  • Qualitative attention visualizations confirm that Xj,kX_{j,k}8 scores highlight salient instances (“9” digits, metastatic regions) and Xj,kX_{j,k}9 scores pinpoint relevant inner-bags.

A plausible implication is that NMIA enhances interpretability for nested weakly-supervised problems and is advantageous where ground-truth is available only at the highest level, but models or applications demand finer-grained insight into hierarchical structure.

NMIA generalizes attention-based MIL architectures via explicit hierarchical nesting, combining soft attention for instance selection with multi-level aggregation. This approach is especially pertinent for domains such as computational pathology, vision, and any application where entities are naturally grouped and only coarse labels are available. The nesting and attention extensibility allow NMIA to subsume previous MIL variants (mean aggregation, single-level attention) and outperform them in tasks necessitating hierarchical inference (Fuster et al., 2021).

The framework's broad applicability suggests future directions in further hierarchy modeling, explainable machine learning, and adaptation to domains with complex nested-label structures.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Nested Multiple Instance Learning with Attention (NMIA).