---
title: Multi-Order Interaction Aggregation
url: https://www.emergentmind.com/topics/multi-order-interaction-aggregation
type: topic
---

# Multi-Order Interaction Aggregation

Searching arXiv for recent papers on multi-order interaction aggregation and closely related terminology.
tool call: arxiv_search({"query":"all:(\"multi-order interaction\" OR \"higher-order\" OR hypergraph) AND (aggregation OR block structure OR filtering)", "max_results": 10, "sort_by": "relevance"})
Multi-Order Interaction Aggregation denotes a family of modeling, inference, and analysis strategies that treat interaction order as an explicit structural variable rather than collapsing all interactions into a single undifferentiated object. In higher-order networks this usually means distinguishing hyperedges of different sizes; in interaction-based interpretability it means stratifying cooperative effects by contextual complexity; in collective decision rules it can mean separating zero-order vote counts from first-order reliabilities and second-order correlations; and in order-theoretic belief change it means aggregating multiple epistemic orderings into a single revised ordering. Across these settings, the common premise is that aggregating all interactions under one universal rule can blur order-dependent mechanisms, whereas order-aware aggregation can yield sharper structure, better prediction, or more principled dynamics [2511.21350, 2305.06910, 2512.18607, 2510.01499].

## 1. Domain-general concept and meanings of “order”

The term *order* is not uniform across the literature. In hypergraphs and higher-order networks, interaction order is the hyperedge size \(|e|\): order \(2\) denotes pairwise interactions, order \(3\) triadic interactions, and so on up to a maximum order \(D\) [2511.21350]. In size-filtered higher-order data analysis, the same scale is described as interaction size, with a \(k\)-hyperedge defined by \(|e|=k\), and a \(k\)-uniform hypergraph containing only \(k\)-hyperedges [2305.06910]. In multi-order Shapley interaction analysis for deep networks, the order \(m\) is not the number of variables in the interacting pair but the size of the contextual subset \(S\) used to evaluate that pair’s joint utility [2512.18607]. In many-body agent systems, the order \(m\) is the number of other agents participating in a simultaneous interaction kernel \(G^{(m)}\) [2502.09098]. In LLM aggregation, “zero-order,” “first-order,” and “second-order” refer respectively to raw answer counts, model-specific accuracies, and pairwise conditional relationships among model outputs [2510.01499].

| Domain | Meaning of order | Aggregation target |
|---|---|---|
| Higher-order networks | Hyperedge size \(|e|\) | Order-specific block structure or filtered hypergraphs |
| DNN interaction analysis | Context size \(m\) for a variable pair | Order-wise interaction components \(\bm{I}^{(m)}\) |
| Multi-agent systems | Number of simultaneously interacting agents | Mean-field operator \(\mathcal{X}^{[m]}\) |
| LLM ensembles | Vote counts, accuracies, pairwise dependencies | OW and ISP decision rules |
| Belief change | Multiple epistemic orderings | TeamQueue-aggregated TPO |

This variation in meaning is central rather than incidental. A persistent misconception is that multi-order aggregation always refers to interaction cardinality. The literature instead uses *order* as a domain-specific index for contextual scale, structural range, or epistemic input multiplicity. What unifies these uses is the refusal to assume that all interactions can be summarized by a single global mechanism.

## 2. Hypergraph disaggregation and size-dependent structure

A major line of work treats multi-order interaction aggregation as a problem of whether higher-order data should be analyzed as one global hypergraph or decomposed by interaction size. A hypergraph is written as \(H=(V,E)\), where \(V\) is the node set and each hyperedge \(e\in E\) is a nonempty subset of \(V\) [2305.06910]. The argument for disaggregation is that small interactions and large interactions may encode different roles or mechanisms, a node can be central in one size regime and unimportant in another, communities may shift when larger groups are included, and global measures may change with interaction size in non-monotone ways [2305.06910]. This motivates the claim that higher-order datasets should not always be treated as one undifferentiated hypergraph.

The filtering framework formalizes this intuition by constructing filtered hypergraphs
\[
H_{f,g}=(V_f,E_g),
\qquad
V_f=\{v\in V\mid f(v)=1\},
\qquad
E_g=\{e\in E\mid g(e)=1\}.
\]
In the size-based setting, the filter is
\[
g(e,k)=
\begin{cases}
1 & \text{if } |e| * k\\
0 & \text{otherwise},
\end{cases}
\]
which yields \(H_{(*,k)}\) [2305.06910]. The main operators are uniform filtering \(H_{(=,k)}\), GEQ filtering \(H_{(\geq,k)}\), LEQ filtering \(H_{(\leq,k)}\), and exclusion filtering \(H_{(\neq,m)}\). Uniform filtering isolates one exact interaction size; LEQ accumulates smaller interactions; GEQ isolates progressively larger interactions; and exclusion filtering measures the sensitivity of structure to one omitted size [2305.06910].

This framework is descriptive rather than generative. It does not posit a latent model for how different orders arise; instead, it creates a family of ordered datasets on which one can compare effective information, assortativity, betweenness centrality, and community labels. On six empirical datasets across email, biology, and proximity networks, the reported patterns include size-dependent changes in effective information, shifts in assortativity peaks, filter-sensitive centrality rankings, and domain-specific stability or instability of communities [2305.06910]. The stated limitations are equally important: filtering reduces the number of interactions, sparse filtered datasets can lead to unstable measures, some changes may reflect sparsity artifacts, there is not yet a theory for choosing when and how to filter, and the paper does not fully quantify “information gain” from considering multiple filterings [2305.06910].

## 3. Multi-order stochastic block structure in higher-order networks

A more explicitly inferential treatment appears in the hypergraph stochastic block-model literature. In the higher-order network setting, a hypergraph is written as \(\mathcal{H}=(\mathcal{V},\mathcal{E})\), the order set is
\[
\mathcal{O}=\{2,3,\ldots,D\},
\]
and the observed data are encoded by \(\mathcal{A}=(A_e)_{e\in\Omega}\), where \(A_e\) counts how many times hyperedge \(e\) appears [2511.21350]. The baseline single-order hypergraph mixed-membership SBM uses a soft membership matrix \(\mathbf{U}\in\mathbb{R}_{+}^{N\times K}\) and a single symmetric affinity matrix \(\mathbf{W}\in\mathbb{R}_{+}^{K\times K}\). Its limitation is that one shared affinity pattern governs all hyperedge sizes [2511.21350].

The multi-order extension, HyperMOSBM, relaxes this assumption by partitioning \(\mathcal{O}\) into disjoint subsets
\[
\mathcal{P}=\{\mathcal{S}_1,\mathcal{S}_2,\ldots,\mathcal{S}_L\},
\]
where each subset \(\mathcal{S}_l\) shares an affinity matrix \(\mathbf{W}^{(l)}\) [2511.21350]. If \(|e|\in\mathcal{S}_{l_{|e|}}\), then the Poisson rate of hyperedge \(e\) uses \(\mathbf{W}^{(l_{|e|})}\), with parameters
\[
\boldsymbol{\theta}=\left(\mathbf{U},\{\mathbf{W}^{(l)}\}_{l=1}^{L}\right).
\]
The generative model is
\[
P(A_e\mid \mathbf{U},\mathbf{W}^{(l_{|e|})})
=
\exp\!\left(-\frac{\lambda_e}{\kappa_{|e|}}\right)
\frac{\left(\frac{\lambda_e}{\kappa_{|e|}}\right)^{A_e}}{A_e!},
\]
with
\[
\lambda_e
=
\sum_{v_i\in e}\sum_{v_j\in e\setminus\{v_i\}}
\sum_{k=1}^{K}\sum_{q=1}^{K}
u_{ik}u_{jq}w_{kq}^{(l_{|e|})},
\qquad
\kappa_s=\binom{s}{2}\binom{N-2}{s-2},
\]
and factorized likelihood
\[
P(\mathcal{A}\mid\boldsymbol{\theta})
=
\prod_{e\in\Omega}P(A_e\mid \mathbf{U},\mathbf{W}^{(l_{|e|})}).
\]
The single-order model is recovered by the trivial partition \(\mathcal{P}=\{\mathcal{O}\}\), \(L=1\); the opposite extreme is the full-order model, which assigns an independent affinity object to each order but is described as much more expensive and liable to overfit [2511.21350].

Partition selection is explicitly out-of-sample. Rather than maximizing in-sample likelihood, the method chooses the partition that maximizes hyperlink prediction AUC under 10-fold cross-validation. For a trained parameter set \(\hat{\boldsymbol{\theta}}\), the hyperlink score is
\[
f_{\hat{\boldsymbol{\theta}}}(e)=\log \hat{\lambda}_e-\log \kappa_{|e|},
\]
which is monotone in \(P(A_e>0)=1-\exp(-\lambda_e/\kappa_{|e|})\) [2511.21350]. The final objective is
\[
\mathcal{P}_{\text{final}}=\arg\max_{\mathcal{P}}\text{AUC}(\mathcal{P}),
\qquad
\Delta_{\text{AUC}}
=
\text{AUC}(\mathcal{P}_{\text{final}})
-
\text{AUC}(\mathcal{P}_1),
\]
where \(\mathcal{P}_1=\{\mathcal{O}\}\) is the single-order baseline. Inference uses EM with variational variables \(\rho_{ijkq}^{(e)}\), an ELBO, E-step updates
\[
\rho_{ijkq}^{(e)}=\frac{u_{ik}u_{jq}w_{kq}^{(l_{|e|})}}{\lambda_e},
\]
and multiplicative M-step updates for \(\mathbf{U}\) and \(\mathbf{W}^{(l)}\) [2511.21350]. For \(|\mathcal{O}|\le 4\), all partitions are tested; otherwise the search is greedy, beginning from \(\{\mathcal{O}\}\) and splitting one subset into two adjacent blocks if AUC improves. The method enforces
\[
\sum_{s\in\mathcal{S}_l} m_s \ge c\,\frac{K(K+1)}{2},
\qquad
c=5,
\]
and stops when no split improves AUC by more than \(10^{-3}\) [2511.21350].

Empirically, across 14 real-world hypergraphs the framework selected a nontrivial partition in 12 of 14 datasets; in almost all of those cases, the AUC gain exceeded \(\Delta_{\text{AUC}}\ge 0.01\); in 9 datasets the gain was statistically significant after Bonferroni correction; and all five co-citation networks exhibited significant multi-order block structure [2511.21350]. Reported examples include the high-school contact partition
\[
\{\{2,4,5\},\{3\}\},
\]
which separates triadic interactions from pairwise, 4-way, and 5-way ones, and co-citation networks in which pairwise co-citations separate from all higher-order ones [2511.21350]. The paper further states that the full-order model was often computationally infeasible and, when it converged, usually performed worse than the multi-order model, indicating overfitting and poor generalization.

## 4. Representation learning and neural interaction structure

In neural representation learning, multi-order interaction aggregation appears as a way of decomposing or constructing model capacity across interaction scales. In the DNN interaction literature, the starting point is the Shapley bivariate interaction index
\[
I(i,j)=\mathbb{E}_m\mathbb{E}_{S\subseteq N\setminus\{i,j\},\,|S|=m}[\Delta v(i,j,S)],
\]
with order-specific term
\[
I^{(m)}(i,j)=\mathbb{E}_{S\subseteq N\setminus\{i,j\},\,|S|=m}[\Delta v(i,j,S)].
\]
The total interaction is then aggregated across orders by
\[
I(i,j)=\mathbb{E}_m[I^{(m)}(i,j)],
\]
and the network output is decomposed as
\[
v(N)=v(\emptyset)+\sum_{i\in N}\mu_i+\sum_{m=0}^{n-2} w^{(m)}\cdot \bm{I}^{(m)},
\]
where \(\bm{I}^{(m)}=\sum_{i,j\subseteq N}I^{(m)}(i,j)\) [2512.18607]. Sample-wise and model-level order strength are summarized by \(S^{(m)}(x)\) and \(J^{(m)}\). Across CNNs, vision transformers, point-cloud networks, NLP transformers, and tabular MLPs, the reported profile of \(J^{(m)}\) is U-shaped: high for small \(m\), low in the middle, and high again for large \(m\) [2512.18607]. The theoretical explanation is that the number of contexts
\[
u(m)=\binom{n-2}{m}
\]
peaks near the middle, yielding a learning strength
\[
F^{(m)}\propto \frac{1}{\sqrt{u(m)}},
\]
so mid-order interactions are hardest to learn. The paper further reports that low-order-emphasized models exhibit stronger generalization and robustness, whereas high-order-emphasized models demonstrate greater structural modeling and fitting capability [2512.18607].

Graph neural and convolutional architectures often operationalize a related idea more constructively. GraphAIR argues that standard GCN-style models largely implement neighborhood aggregation while capturing neighborhood interactions only weakly through nonlinearities; Proposition 1 states that, under a sigmoid expansion, the coefficient of the high-order interacting terms is at most \(1/48\) [1911.01731]. Its remedy is an explicit interaction module
\[
\mathbf{h}_i^{\mathrm{ir}}
=
\left(\sum_{j \in \mathcal{N}_i} e_{ij}\mathbf{h}_j\mathbf{W}\right)
\odot
\left(\sum_{k \in \mathcal{N}_i} e_{ik}\mathbf{h}_k\mathbf{W}'\right),
\]
combined with an aggregation branch through
\[
\mathbf{h}_i^{\mathrm{air}}=\sigma(\mathbf{h}_i^{\mathrm{agg}})+\sigma(\mathbf{h}_i^{\mathrm{ir}}).
\]
MogaNet similarly frames modern ConvNets as suffering from a representation bottleneck in which extreme-order interactions dominate at the expense of middle-order ones. Its spatial module combines feature decomposition, multi-order depth-wise convolutions, and gated aggregation, with low-, middle-, and high-order branches instantiated as \(\mathrm{DW}_{5\times5,d=1}\), \(\mathrm{DW}_{5\times5,d=2}\), and \(\mathrm{DW}_{7\times7,d=3}\) [2211.03295]. PanCAN extends this logic to multi-label vision by recursively defining \(k\)-th order neighborhoods
\[
{\cal N}_c^{(k)}(\mathbf{x})
=
\bigcup_{\mathbf{x}'\in {\cal N}_c^{(k-1)}(\mathbf{x})}
{\cal N}_c^{(k-1)}(\mathbf{x}'),
\]
then aggregating order-specific features via attention-weighted random walks and cross-scale anchor-based fusion [2512.23486]. CS-IGANet applies an analogous decomposition to mouse social behavior, explicitly aggregating intra-skeleton, inter-skeleton, and cross-skeleton interactions before hierarchical graph-level pooling [2208.03819].

These architectures do not all define *order* identically, but they share a design principle: an additive or multiplicative aggregate over multiple interaction regimes is more expressive than a single undifferentiated neighborhood summary.

## 5. Collective inference, many-body limits, and order aggregation outside statistical learning

Multi-order interaction aggregation also appears in domains where the goal is not representation learning but collective inference or dynamical reduction. In LLM answer aggregation, majority voting is treated as a zero-order rule because it uses only the raw votes \(a_1,\dots,a_N\). The Optimal Weight rule incorporates first-order information, namely model accuracies \(x_i=P(A_i=S^*)\), through
\[
f_{OW}(a_1,\dots,a_N)=\arg\max_{s\in S}\sum_{i=1}^N \sigma_K^{-1}(x_i)\,\ind\{a_i=s\},
\]
while Inverse Surprising Popularity uses second-order information through conditional relations \(P(A_i=s_k\mid A_j=s_l)\) and the counterfactual score
\[
S_{ISP}(s,i)=\frac{1}{N-1}\sum_{j\neq i}\frac{1}{K-1}\sum_{a\in S\setminus\{a_j\}}P(A_i=s\mid A_j=a).
\]
The paper proves that \(f_{OW}\) is Bayesian optimal under conditional independence and reports that, across synthetic data, UltraFeedback, MMLU, and ARMMAN, OW and ISP consistently outperform majority voting [2510.01499]. Here aggregation does not partition orders; it augments decision rules by incorporating progressively richer orders of information.

In interacting multi-agent systems, the order \(m\) indexes simultaneous many-body couplings. The microscopic dynamics are
\[
\dot\xi_i^N(t)
=
\frac{1}{N^m}\sum_{j_1,\dots,j_m=1}^N
G^{(m)}\!\left(t,x_i^N,\xi_i^N(t),x_{j_1}^N,\xi_{j_1}^N(t),\dots,x_{j_m}^N,\xi_{j_m}^N(t)\right),
\]
and the associated mean-field aggregation operator is
\[
\mathcal{X}^{[m]}[f](t,x,\xi)
=
\int G^{(m)}(\cdots)\,d f^{\otimes m}.
\]
This yields the mesoscopic Vlasov-type equation
\[
\partial_t f+\nabla_\xi\cdot\big(\mathcal{X}^{[m]}[f]\,f\big)=0,
\]
together with well-posedness, a Dobrushin-type stability estimate, propagation of chaos, and a large-order limit \(m\to\infty\) in which the many-body interaction collapses to an effective macroscopic law \(\partial_t y^\infty=G^\infty_\nu[y^\infty]\) [2502.09098]. In this setting, aggregation is a rigorous averaging over all \(m\)-tuples rather than a predictive model-selection device.

A formally different but conceptually related use appears in iterated belief change. In parallel belief revision, revision by a package \(S=\{A_1,\dots,A_n\}\) is not treated as mere revision by \(\bigwedge S\) in the iterated case. Instead, one revises separately by each \(A_i\), aggregates the resulting total preorders using a TeamQueue order aggregator \(\oplus\), and then enforces success by a final revision step [2505.13914]. The output ordering is built from profiles \(\mathbf{P}=\langle \preccurlyeq_i\rangle_{i\in I}\) via stages
\[
T_i=\bigcup_{j\in a_{\mathbf{P}(i)}}\min\!\left(\preccurlyeq_j,\bigcap_{k<i}T_k^c\right),
\]
with the synchronous version taking \(a_{\mathbf{P}(i)}=\{1,\dots,n\}\) at every step [2505.13914]. The related theory of parallel contraction generalizes the same TeamQueue logic to \(n\)-ary order aggregation and characterizes the aggregator by the factoring property
\[
\text{For all }S\subseteq W,\ \exists X\subseteq I,\ \text{s.t. }
\min(\preccurlyeq_{\oplus},S)=\bigcup_{j\in X}\min(\preccurlyeq_j,S),
\]
with additional parity and flattest-order results [2501.13295]. In these settings, *interaction aggregation* concerns the combination of multiple revised orderings into a single coherent epistemic order.

## 6. Interpretive value, empirical gains, and recurrent limitations

A central interpretive claim across the literature is that order-aware aggregation sharpens latent structure. In higher-order network modeling, the single-order model tends to blur different kinds of interactions into one broad community-affinity pattern, whereas the multi-order model can produce clearer community boundaries, better recovery of known ground-truth classes, more interpretable affinity matrices, and a sharper mapping between observed hyperedges and latent communities [2511.21350]. In the high-school data, the multi-order fit is reported to align inferred communities almost perfectly with class labels; in the computer-science co-citation network it yields a cleaner and more coherent topical organization [2511.21350]. The filtering literature makes a parallel point in descriptive terms: the whole hypergraph can hide the fact that different scales carry different information, so the “sum of the parts may be greater than the whole” [2305.06910].

At the same time, the literature repeatedly warns against the opposite extreme of unconstrained order separation. HyperMOSBM contrasts its partition-based approach with the full-order model, which was often computationally infeasible and, when it converged, usually performed worse, indicating overfitting and poor generalization [2511.21350]. The filtering framework emphasizes noise sensitivity, instability on sparse filtered graphs, and the lack of a theory for choosing when and how to filter [2305.06910]. In DNN interaction analysis, the problem is not over-aggregation but a built-in representational bias: mid-order interactions are intrinsically difficult to learn because contextual variability is maximal there [2512.18607]. In belief revision, overly strong postulates such as SC2 are explicitly rejected; TeamQueue parallel revision operators recover Conj, PC3, PC4, C1–C4, Ind, and S under the stated assumptions, but not SC2 or P [2505.13914]. A plausible implication is that multi-order aggregation is most useful when it is neither collapsed to a single universal rule nor expanded to a maximally unconstrained per-order model.

Taken together, these results define multi-order interaction aggregation as an organizing principle rather than a single method. Its recurring task is to determine when heterogeneous interaction scales should be kept separate, when they should share parameters, and how their contributions should be recombined. In hypergraphs this leads to order-partitioned block structure; in neural networks it leads to order-wise decomposition or explicit interaction modules; in ensemble decision-making it yields higher-order voting rules; in many-body systems it produces mean-field and macroscopic limits; and in belief change it yields principled aggregation of epistemic orderings. The shared conclusion is not that finer order resolution is always preferable, but that the structure of interactions often depends on order, and that any aggregation rule ignoring this dependence risks obscuring the very phenomena it is meant to summarize [2511.21350, 2305.06910].

Source: https://www.emergentmind.com/topics/multi-order-interaction-aggregation