---
title: Submodular Information Measures (SIM)
url: https://www.emergentmind.com/topics/submodular-information-measures-sim
type: topic
---

# Submodular Information Measures (SIM)

A submodular information measure (SIM) is a combinatorial generalization of classical information-theoretic measures, such as entropy, mutual information, and conditional mutual information, in which the foundational role of Shannon entropy is replaced with a general monotone submodular set function. SIMs provide an abstract algebraic framework for modeling information, relevance, coverage, independence, and diversity on arbitrary ground sets, extending beyond random variables to structured data, feature sets, and combinatorial objects. SIMs have deep implications across data subset selection, active learning, summarization, privacy, representation learning, causal inference, and extremal combinatorics.

## 1. Formal Definitions and Mathematical Structure

Let $V$ be a finite ground set and $f: 2^V \to \mathbb{R}$ a normalized, monotone, submodular set function: $f(\emptyset) = 0$, $f(A) \leq f(B)$ whenever $A \subseteq B$, and for all $A, B \subseteq V$, $f(A) + f(B) \geq f(A \cup B) + f(A \cap B)$.

The key submodular information measures are:

- **Submodular Conditional Gain (SCG):**
  $$ f(A \mid P) := f(A \cup P) - f(P) $$
  Intuition: The incremental "utility" provided by $A$ beyond $P$.

- **Submodular Mutual Information (SMI):**
  $$ I_f(A; Q) := f(A) + f(Q) - f(A \cup Q) $$
  Intuition: The amount of “shared information” or representativeness of $A$ with respect to $Q$.

- **Submodular Conditional Mutual Information (SCMI):**
  $$ I_f(A; Q \mid P) := f(A \cup P) + f(Q \cup P) - f(A \cup Q \cup P) - f(P) $$
  Equivalently, $I_f(A; Q \mid P) = I_f(A; Q \cup P) - I_f(A; P)$. Intuition: Relevance of $A$ to $Q$ penalized by overlap with $P$.

These extend immediately to multi-set analogues, total correlation, and composite objectives. For example, the total correlation of $k$ disjoint sets is $C_f(A_1, ..., A_k) = \sum_{i=1}^k f(A_i) - f(\cup_{i=1}^k A_i)$ [2310.00165].

When $f$ is the entropy of a collection of random variables, these recover classical Shannon-information measures. For canonical submodular functions like coverage, facility-location, concave-over-modular, or certain graph-cut-type objectives, SIMs coincide exactly with entropic mutual information under explicit constructions [2601.12724].

## 2. Theoretical Properties: Axioms and Independence

SIMs inherit critical properties from submodularity [2108.03154, 2006.15412]:

- **Nonnegativity:** $I_f(A; B) \geq 0$ and $f(A|P) \geq 0$ for normalized, monotone $f$.
- **Symmetry:** $I_f(A; B) = I_f(B; A)$.
- **Monotonicity:** $A \mapsto I_f(A; B)$ is non-decreasing for fixed $B$; $f(A|P)$ is monotone in $A$.
- **Submodularity in One Argument:** $A \mapsto I_f(A; B)$ is submodular in $A$ when $f$'s third discrete derivatives are non-negative; this holds for facility-location, set cover, concave-over-modular, and some graph-cut functions [2006.15412, 2105.00043].
- **Chain Rule:** $I_f(A; B \cup C) = I_f(A; C) + I_f(A; B|C)$.

Independence concepts are generalized:
- **Joint Independence:** $A \perp_J B \iff I_f(A; B) = 0$.
- **Pairwise Independence:** $A \perp_P B$ if for all $a \in A, b \in B$, $I_f(\{a\};\{b\}) = 0$.
- **Multi-set Independence:** $C_f(A_1, ..., A_k) = 0$ [2108.03154].

These fundamental axioms enable the use of SIMs in combinatorial optimization, privacy, summarization, and learning tasks that require formal guarantees.

## 3. Canonical Submodular Function Classes and Entropic Correspondence

The most widely used SIMs are grounded in the following classes [2601.12724, 2006.15412]:

| Function family    | $f(A)$ definition                                                         | Typical use cases                  |
|--------------------|--------------------------------------------------------------------------|------------------------------------|
| Coverage/set-cover | $\sum_{u \in \cup_{i\in A} U_i} w_u$                                    | Diversity, coverage                |
| Facility-location  | $\sum_{i\in V} \max_{j\in A} s_{ij}$                                    | Representation, information overlap|
| Graph-cut-type     | $\sum_{i\in V, j\in A} s_{ij} - \lambda \sum_{i,j\in A}s_{ij}$          | Redundancy, separation, clustering |
| Concave-over-mod   | $\psi(\sum_{j\in A} w_j)$, $\psi$ concave nondecreasing                 | Robustness, budgeted diversity     |
| Log-determinant    | $\log \det(S_A + \epsilon I)$ ($S$ kernel)                              | Volume, diversity, uncertainty     |

Recent work demonstrates exact entropic constructions: given any of these $f$, there exists a random vector $(X_i)_{i \in V}$ so that $f(A) = H(X_A)$, and all submodular information measures reduce to their classical Shannon counterparts [2601.12724].

## 4. Optimization Algorithms and Greedy Guarantees

Maximization of any nonnegative, monotone SIM (e.g., $I_f(A; Q)$, $f(A|P)$, $I_f(A; Q|P)$) under a cardinality or matroid constraint admits a $(1-1/e)$-approximation via the greedy algorithm [2206.08566, 2105.00043, 2103.00128]:

1. Initialize $A \gets \emptyset$.
2. For $i=1$ to $B$: 
   - For each $u$ not in $A$, compute marginal gain: e.g., $\Delta(u) = I_f(A\cup\{u\}; Q) - I_f(A; Q)$.
   - Add $u^*$ with maximal $\Delta(u)$ to $A$.

Lazy-greedy and partitioning reduce computational cost, particularly for SMI based on facility-location (requiring only $|Q|\times|U|$ similarity evaluations) [2206.08566, 2105.00043].

Curvature bounds ([curvature $\kappa_f$]) further tighten approximation ratios. In practice, facility-location, graph-cut, and log-determinant functions exhibit low curvature, making greedy nearly optimal [2206.08566].

## 5. Applications: Data Selection, Summarization, and Learning

SIMs constitute core objectives in broad machine learning settings:

- **Active Learning and Data Discovery:** SCG and SMI are used to mine rare or unknown classes by rewarding dissimilarity from labeled sets (SCG) and then intensifying discovery by targeting known hits (SMI/SCMI). Empirically, these approaches dominate baselines on rare-class and OOD selection in image classification and object detection, with 10–15% absolute gains in accuracy for unknowns [2206.08566, 2107.00717, 2210.01526].

- **Targeted Subset Selection:** SMI and variants (facility-location, log-det, graph-cut, COM) select samples that optimally trade off query relevance and target coverage. Theoretical bounds guarantee that maximizing SMI under realistic similarity-separation assumptions ensures high query relevance and coverage [2402.13454, 2105.00043].

- **Privacy and Fairness:** SCMI and its constraints operationalize privacy by enforcing independence from a sensitive set under a user-defined threshold [2108.03154, 2010.05631]. Privacy filters and marginal-independence filters compose efficiently with submodular maximization objectives.

- **Summarization and Representation Learning:** SIMs unify generic, query-focused, privacy- and update-aware summarization as direct maximizations of SMI, SCG, or CSMI, generalizing models such as ROUGE, DPPs, and graph-cut methods [2010.05631, 2103.00128]. In representation learning, submodular total correlation losses (e.g., SCoRe framework) simultaneously minimize intra-class variance and inter-class bias, outperforming standard contrastive methods for imbalanced data [2310.00165].

A selection of empirical results:

| Application | SIM Instantiation   | Typical Gain over Baselines      | Reference  |
|-------------|---------------------|----------------------------------|------------|
| Active Data Discovery | Fl_cg+mi, Logdet_cg+mi | 10–15% higher accuracy on unknowns | [2206.08566] |
| OOD Avoidance | Fl-CMI, LogDet-CMI | 4–7% accuracy lift | [2210.01526] |
| Targeted TSS | LogdetMI, FL2MI    | ~20–30% absolute improvement on rare classes | [2105.00043] |
| Summarization | FL-SMI, GraphCut-SMI, LogDet-SMI | Near human-level V-ROUGE | [2010.05631] |
| Representation Learning | FL-/GC -C_f | 1–9% boost in class-imblanced recognition | [2310.00165] |

## 6. Extensions: Causal Inference, Information Inequalities, and Advanced Properties

SIMs extend classical independence, conditional independence, and causal Markov properties to non-entropic settings, unifying information-theoretic and combinatorial perspectives. The generalized causal Markov condition for SIMs matches the standard DAG-based independence structure, independent of the choice of submodular $f$ [1002.4020].

Unified derivations of information inequalities (Han’s, Shearer’s, monotonicity sequences, total correlation bounds) follow broadly from submodularity. These yield refined combinatorial bounds, e.g., on projection sizes, Boolean influences, and extremal graph properties [2204.13410].

Recent developments include the study of SIMs for weak submodularity in quadratic estimation and optimal experimental design (alphabetic optimality criteria), where closed-form utility functions (log-det, trace, min-eigenvalue) are directly submodular or enjoy quantifiable approximation via greedy [1905.09919].

## 7. Modeling Flexibility, Parameterizations, and Practical Considerations

Modern extensions, such as PRISM [2103.00128], introduce multi-parameterized SIMs to interpolate between relevance, diversity, privacy, and coverage. Typical parameters:
- $\lambda$ (graph-cut): relevance vs. diversity
- $\eta$ (facility-location, COM): similarity-to-query trade-off
- $\nu$ (conditional gain): strength of avoidance/penalty to a private set

By tuning these, SIMs adapt to a wide regime of problems: rare-class mining, guided summarization, OOD filtering, distributed and scalable optimization.

Several concrete choices are supported with efficient greedy algorithms [2103.00128, 2010.05631, 2402.13454], and the entire framework is modality-agnostic—applicable to images, video, text, sensor sets, and gradient embeddings.

---

**Summary Table of Core SIM Formulae**

| Name        | Formula                                                    | Typical Use                   |
|-------------|------------------------------------------------------------|-------------------------------|
| SCG         | $f(A|P) = f(A\cup P) - f(P)$                              | Dissimilarity, novelty        |
| SMI         | $I_f(A;Q) = f(A) + f(Q) - f(A \cup Q)$                    | Relevance, coverage, overlap  |
| SCMI        | $I_f(A;Q|P) = f(A \cup P) + f(Q \cup P) - f(A \cup Q \cup P) - f(P)$ | Targeting under exclusion     |

---

Through their algebraic generality and foundational approximation guarantees, submodular information measures constitute a principled, tractable, and highly expressive toolkit for information-centric decision-making in structured data systems [2206.08566, 2108.03154, 2006.15412, 2402.13454, 2105.00043].

Source: https://www.emergentmind.com/topics/submodular-information-measures-sim