Papers
Topics
Authors
Recent
Search
2000 character limit reached

Aggregate Mutual Information in Frequency (AMIF)

Updated 14 July 2026
  • AMIF is an information-theoretic measure that aggregates frequency-resolved mutual information between stochastic processes to capture both same-frequency and cross-frequency dependencies.
  • It is computed by estimating mutual information via FFT-based spectral increments and applying nonparametric k-nearest-neighbor methods under stationarity assumptions.
  • AMIF underpins visualization and clustering in O-RAN KPI analysis by converting dependence measures into dissimilarities for effective core KPI selection.

Searching arXiv for the cited AMIF and MI-in-frequency papers to ground the article in current records. Aggregate Mutual Information in Frequency (AMIF) is an information-theoretic construction that aggregates frequency-resolved mutual information between stochastic processes. It is rooted in mutual information in frequency (MIF), which measures dependence between spectral process increments in Cramér’s representation rather than between raw time-domain samples, and is therefore designed for dependent time-series rather than unordered observations (Malladi et al., 2017). In Gaussian linear time-invariant settings, AMIF reduces to an integral of same-frequency mutual information and coincides with the mutual information rate per sample; in more general nonlinear or non-Gaussian settings, the literature uses aggregation over statistically significant coupled frequencies to capture multi-frequency dependence without simple diagonal additivity (Malladi et al., 2017). In O-RAN testing, AMIF is the central similarity measure in the AMIF-MDS workflow for visualization, clustering, and “core” KPI selection from PHY/MAC-layer measurements (Pradhan et al., 3 Oct 2025). The acronym is not entirely uniform across domains: related arXiv usages include a frequency-domain computation of mutual information over all image translations and an operator-theoretic frequency aggregation for frequency-selective MIMO channels (Öfverstedt et al., 2021, Hachem et al., 2015).

1. Conceptual and mathematical foundations

The foundational idea behind AMIF is that temporal dependence should be analyzed in the spectral domain. For a zero-mean, second-order stationary discrete-time process, the Cramér representation writes the process as an integral over orthogonal spectral increments. In the notation used in the MIF literature,

X[n]=1/21/2ei2πfndZX(f),Y[n]=1/21/2ei2πfndZY(f).X[n] = \int_{-1/2}^{1/2} e^{i 2\pi f n} dZ_X(f), \qquad Y[n] = \int_{-1/2}^{1/2} e^{i 2\pi f n} dZ_Y(f).

MIF is then defined as the mutual information between the two-dimensional real vectors formed by the real and imaginary parts of spectral increments at selected frequencies:

MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).

This formulation permits dependence at the same frequency and across different frequencies, thereby covering cross-frequency coupling as well as conventional same-band dependence (Malladi et al., 2017).

For jointly Gaussian stationary processes, the frequency-resolved quantity collapses to a coherence-based expression:

IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),

where γ2(f)\gamma^2(f) is the magnitude-squared coherence. In this special case, AMIF is the integral of same-frequency MIF:

AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,

or, on a discrete grid,

I^(X;Y)=1Nfi=0Nf/2MI^XY(λi;λi).\hat I(X;Y)=\frac{1}{N_f}\sum_{i=0}^{N_f/2}\widehat{MI}_{XY}(\lambda_i;\lambda_i).

The units are nats per sample when natural logarithms are used, and bits per sample after division by ln2\ln 2 (Malladi et al., 2017).

For nonlinear or non-Gaussian coupling, additive integration along the diagonal is generally insufficient because distinct frequency components can be statistically dependent. The nonparametric aggregation proposed in the original MIF estimator forms sets Λx\Lambda_x and Λy\Lambda_y of frequencies participating in statistically significant dependence and estimates

I^(X;Y)=1max(P,Q)I^(dX~(Λx);dY~(Λy)),\hat I(X;Y)=\frac{1}{\max(P,Q)}\hat I\big(d\widetilde{X}(\Lambda_x);d\widetilde{Y}(\Lambda_y)\big),

with MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).0 and MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).1. This is the canonical AMIF construction in the dependent-data literature (Malladi et al., 2017). In the neuroscience formulation, the same framework is used to define bandwise and cross-band aggregates, although the paper does not introduce a single scalar explicitly named AMIF; instead, the integral relation for Gaussian processes and weighted sums over frequency-pair grids supply the principled aggregation rules (Malladi et al., 2017).

2. Estimation from finite data

Finite-sample AMIF estimation proceeds by approximating spectral increments with Fourier coefficients computed on nonoverlapping windows. A time-series of length MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).2 is partitioned into MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).3 windows of length MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).4, where MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).5, and for each window the FFT provides approximate samples of MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).6 and MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).7. Under stationarity and mixing assumptions, these window-level spectral samples behave as approximately i.i.d. observations of the underlying increment variables, making nonparametric MI estimation tractable in low-dimensional spaces (Malladi et al., 2017).

The standard estimator in the MIF literature is the Kraskov–Stögbauer–Grassberger k-nearest-neighbor estimator applied to the four-dimensional real vector composed of the real and imaginary parts of the two spectral increments. With MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).8, the estimator is

MIXY(λi,λj)=I({dX~R(λi),dX~I(λi)};{dY~R(λj),dY~I(λj)}).MI_{XY}(\lambda_i,\lambda_j) = I\big(\{d\widetilde{X}_R(\lambda_i), d\widetilde{X}_I(\lambda_i)\}; \{d\widetilde{Y}_R(\lambda_j), d\widetilde{Y}_I(\lambda_j)\}\big).9

where IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),0 and IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),1 count neighbors in the marginal subspaces and IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),2 is the digamma function. The literature also describes a kernel-density alternative, but the kNN estimator is reported to converge faster, exhibit lower bias, and run faster computationally in the validation studies summarized in the data block (Malladi et al., 2017).

Significance assessment is performed by permutation tests on the window-level spectral samples. For each frequency pair, the samples of one process are randomly permuted across windows, the MI estimate is recomputed IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),3 times, and the observed value is compared with the resulting null distribution. Significant pairs define the coupled-frequency sets used in the general AMIF aggregation, or, in diagonal-only settings, justify diagonal summation (Malladi et al., 2017).

The O-RAN AMIF estimator modifies this earlier permutation-based approach. Rather than testing all frequency pairs by permutation, it “replace[s] the original permutation testing with a highly efficient quantile-based selection of significant frequency pairs.” Operationally, the estimator segments the two series, applies FFT to each segment, computes MI for all frequency-bin pairs using a k-NN estimator on the complex FFT samples, selects the top-IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),4 fraction of scores, concatenates the selected spectral samples within each series, computes one final MI between the two aggregate matrices, and then normalizes the result to obtain the AMIF score. The paper emphasizes robustness through frequency-domain analysis, nonparametric k-NN MI estimation, quantile filtering, and aggregation across selected frequencies, but it does not print an explicit closed-form equation for IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),5 or for the final AMIF functional (Pradhan et al., 3 Oct 2025).

3. AMIF-MDS in O-RAN KPI analysis

In “Mutual Information-Driven Visualization and Clustering for Core KPI Selection in O-RAN Testing” (Pradhan et al., 3 Oct 2025), AMIF is used as the similarity measure in a complete pipeline for KPI analysis. The paper motivates the problem by the rapid growth in O-RAN performance measurements as systems become more complex, with additional units, interfaces, applications, implementations, and configurations. Because these KPIs are time-series and may exhibit nonlinear dependencies, ordinary sample-level mutual information is treated as inadequate, while direct directed-information or transfer-entropy estimation is regarded as difficult on continuous real-world data.

The AMIF-MDS workflow begins with pairwise AMIF estimation between all KPI time-series. After the similarity matrix is assembled, it is post-processed in two specific ways: it is made perfectly symmetric by averaging it with its transpose, and its diagonal elements are set to infinity because “the mutual information between a continuous random variable and itself is theoretically infinite” (Pradhan et al., 3 Oct 2025). The resulting matrix is converted into a dissimilarity matrix by one of two transformations. The membership transformation is

IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),6

and the logarithmic transformation is

IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),7

These dissimilarities are embedded in a low-dimensional Euclidean space by classical multidimensional scaling. In the O-RAN study, the visualization is three-dimensional, with shadows on the bottom plane to aid interpretation (Pradhan et al., 3 Oct 2025).

Clustering is then applied to the MDS embedding using DBSCAN. The reported parameters are “a neighborhood radius of 0.15 and a minimum cluster size of 1.” No formal cluster-validity index is reported; interpretation is instead tied to domain knowledge and spatial grouping in the embedding (Pradhan et al., 3 Oct 2025). This pairing of AMIF, MDS, and DBSCAN converts a pairwise dependence matrix into an interpretable geometric representation of mutually informative KPIs.

4. Synthetic validation and O-RAN findings

The O-RAN paper validates the method on a synthetic benchmark built from eight independent AR(3) parent processes with random linear trends and nonlinear children formed by squaring each parent. The generation step is

IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),8

to which a trend IXY(f)=12log(1γ2(f)),I_{XY}(f) = -\tfrac{1}{2}\log\big(1-\gamma^2(f)\big),9 is added with γ2(f)\gamma^2(f)0 and γ2(f)\gamma^2(f)1, after which γ2(f)\gamma^2(f)2 is formed element-wise and each series is normalized to zero mean and unit variance (Pradhan et al., 3 Oct 2025). On this dataset, Euclidean distance, maximum absolute cross-correlation, and maximum absolute correlation coefficient fail to reveal the true parent-child structure because of nonlinearities and trends, whereas AMIF-based dissimilarities “consistently exhibit a pronounced block-diagonal structure, accurately identifying all eight parent-child clusters irrespective of the tested AMIF parameters or transformation type.” The reported parameter ranges are γ2(f)\gamma^2(f)3 and γ2(f)\gamma^2(f)4, with both membership and logarithmic dissimilarities (Pradhan et al., 3 Oct 2025).

The applied O-RAN study uses thirteen PHY/MAC-layer KPIs sampled every γ2(f)\gamma^2(f)5 ms under random OFDM burst interference, with Packet Delay excluded because of incompleteness. AMIF similarity, membership transformation, classical MDS in three dimensions, and DBSCAN with γ2(f)\gamma^2(f)6 and minimum cluster size γ2(f)\gamma^2(f)7 produce a structured KPI map in which the largest cluster corresponds to the downlink link-adaptation chain: MAC-DL-CQI, DL-SINR, RSRP/RSRQ, and PHY-MCS. The reported interpretation is that physical-layer measurements such as RSRP, RSRQ, and DL-SINR inform the user-reported CQI, which in turn governs PHY-MCS selection, explaining the strong mutual information and co-clustering (Pradhan et al., 3 Oct 2025).

Additional cluster structure refines the KPI landscape. Spectral efficiency is a singleton positioned near PHY-MCS and MAC-N-PRB, consistent with the statement that “SE is shaped by both… PHY-MCS and scheduling parameter MAC-N-PRB.” RSSI clusters with MAC-DL-PMI, which the authors interpret as interference peaks synchronized with spatial-processing updates under burst jamming. Other singleton clusters include MAC-DL-RI, MAC-UL-Buffer, MAC-N-PRB, and DL-BLER, each treated as contributing largely orthogonal information or occupying a distinct position relative to the link-adaptation cluster (Pradhan et al., 3 Oct 2025). The practical outcome is a “core KPI set” centered on link-adaptation indicators and closely related measures, intended to support streamlined testing and feature selection.

The term AMIF is not used identically across all arXiv contexts. In multimodal image alignment, “Fast computation of mutual information in the frequency domain with applications to global multimodal image alignment” interprets frequency-domain MI aggregation as the cross-mutual information function (CMIF) evaluated over all discrete translations. Here the core object is

γ2(f)\gamma^2(f)8

where the joint histogram counts at shift γ2(f)\gamma^2(f)9 are cross-correlations of indicator images. The frequency-domain acceleration comes from FFT-based evaluation of these correlations:

AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,0

In this setting, AMIF is effectively a dense MI map over translations, aggregated by operations such as AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,1 to recover the best alignment. The dominant complexity is reduced from a direct spatial method scaling as AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,2 to an FFT-based method scaling as AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,3, and reported GPU speed-ups range from about AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,4 to more than AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,5 for realistic image sizes (Öfverstedt et al., 2021).

A different usage appears in the operator-theoretic study of frequency-selective MIMO channels. There, AMIF denotes mutual information per receive antenna aggregated across the spectrum of an ergodic self-adjoint operator associated with the channel. If AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,6 is the Integrated Density of States of AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,7, the per-antenna mutual information is

AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,8

and, with explicit SNR scaling,

AMIF=01/2IXY(f)df,\mathrm{AMIF} = \int_{0}^{1/2} I_{XY}(f)\,df,9

This usage aggregates across both frequency selectivity and spatial modes and is explicitly distinguished from “average mutual information” in time-delay estimation contexts (Hachem et al., 2015).

This diversity of meanings does not invalidate the term, but it makes context essential. In the dependent-data and O-RAN literature, AMIF refers to aggregation of mutual information in frequency between stochastic processes. In image alignment, it is a frequency-domain strategy for evaluating MI over all discrete displacements. In wideband MIMO, it is a spectral integral over the channel operator’s density of states. A plausible implication is that citation context is necessary whenever AMIF is introduced outside a narrowly defined subfield.

6. Assumptions, limitations, and recurrent misconceptions

A recurrent misconception is that AMIF is inherently directional. The O-RAN paper positions AMIF as a practical proxy for directed information because, under certain conditions, aggregated MIF can correspond to mutual information rate or even DI. However, the estimator used there is symmetric, the similarity matrix is explicitly symmetrized, and no directionality inference method is proposed in that work (Pradhan et al., 3 Oct 2025). AMIF therefore quantifies dependence strength in that application, not causal direction.

Another misconception is that AMIF has a single universal formula. The literature instead presents a family of related constructions. In the Gaussian LTI case, the diagonal integral

I^(X;Y)=1Nfi=0Nf/2MI^XY(λi;λi).\hat I(X;Y)=\frac{1}{N_f}\sum_{i=0}^{N_f/2}\widehat{MI}_{XY}(\lambda_i;\lambda_i).0

is exact. For nonlinear or non-Gaussian systems with cross-frequency coupling, the general aggregation is joint over the significant frequency sets rather than a simple sum along the diagonal. The neuroscience paper explicitly notes that it does not define a single scalar called AMIF, even though bandwise aggregation consistent with the MIF framework is natural (Malladi et al., 2017).

Methodologically, the time-series formulations rely on stationarity or local stationarity within windows, sufficient segment counts for nonparametric MI estimation, and FFT-based approximations to spectral increments. The original MIF papers discuss spectral leakage, window-length selection, and finite-sample effects, while the O-RAN paper notes sensitivity to the number of frequency bins I^(X;Y)=1Nfi=0Nf/2MI^XY(λi;λi).\hat I(X;Y)=\frac{1}{N_f}\sum_{i=0}^{N_f/2}\widehat{MI}_{XY}(\lambda_i;\lambda_i).1, the quantile I^(X;Y)=1Nfi=0Nf/2MI^XY(λi;λi).\hat I(X;Y)=\frac{1}{N_f}\sum_{i=0}^{N_f/2}\widehat{MI}_{XY}(\lambda_i;\lambda_i).2, segment length, overlap, and the choice of k-NN estimator, without providing a formal sensitivity analysis or explicit complexity bounds (Malladi et al., 2017, Pradhan et al., 3 Oct 2025). The O-RAN study also does not report small-sample bias corrections, multitaper or Welch smoothing, or formal cluster-validity indices for DBSCAN (Pradhan et al., 3 Oct 2025).

Within-process matrices require additional caution because the diagonal term is theoretically infinite for a continuous random vector’s mutual information with itself. This appears both in the general MIF literature, where I^(X;Y)=1Nfi=0Nf/2MI^XY(λi;λi).\hat I(X;Y)=\frac{1}{N_f}\sum_{i=0}^{N_f/2}\widehat{MI}_{XY}(\lambda_i;\lambda_i).3, and in the O-RAN implementation, where the diagonal of the similarity matrix is set to infinity before dissimilarity transformation (Malladi et al., 2017, Pradhan et al., 3 Oct 2025). Interpretation should therefore focus on cross-process dependence or off-diagonal structure, depending on the application.

Taken together, the arXiv literature presents AMIF as a technically flexible but context-sensitive construct. In its core time-series sense, it is a frequency-domain aggregation of mutual information designed to estimate dependence between random processes under temporal structure, with special simplifications in Gaussian LTI settings and nonparametric joint aggregation in nonlinear or non-Gaussian cases (Malladi et al., 2017). Its recent O-RAN instantiation demonstrates how such a measure can drive visualization and clustering for KPI reduction, while adjacent domains show that the same acronym can legitimately denote different forms of frequency-domain mutual-information aggregation (Pradhan et al., 3 Oct 2025, Öfverstedt et al., 2021, Hachem et al., 2015).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Aggregate Mutual Information in Frequency (AMIF).