Relational Hyper-Event Models (RHEMs)
- RHEMs are advanced event-history models that generalize dyadic interactions to hyperedges, capturing group-level and higher-order dependencies.
- They assign event rates to candidate hyperedges using Cox-type proportional hazards models with covariates derived from past events and exogenous factors.
- RHEMs efficiently model complex events like coauthorship, multicast communications, and scientific collaborations via scalable estimation and sampled risk sets.
Relational Hyper-Event Models (RHEMs) are event-history models that generalize relational event models from dyadic events to events on hyperedges: arbitrary subsets of actors in undirected settings, or pairs of source and target subsets in directed settings. In this framework, the basic stochastic unit is no longer a pair , but a set-valued or set-pair-valued interaction such as a meeting, a coauthor team, a multicast communication, or a publication linking authors to references. RHEMs assign event rates to candidate hyperedges in a risk set and explain realized events through hyperedge covariates derived from past events and exogenous information. This makes them suitable for settings in which decomposing one event into multiple dyads would impose invalid dyadic independence, lose information about group composition, or obscure higher-order dependence (Lerner et al., 2019, Lerner et al., 2021).
1. Formal object and model class
A foundational formulation starts from hypergraphs. In an undirected hypergraph,
where is a finite node set and is a set of undirected hyperedges; each undirected hyperedge is a set . In a directed hypergraph,
where ; each directed hyperedge is a pair
with the source set and the target set. A relational hyperevent is then a time-stamped event on such a hyperedge. One common notation is
0
where 1 is the hyperedge, 2 the event time, 3 an event type and/or weight, and 4 a relational outcome; another is 5, with 6 a set of senders and 7 a set of receivers. In undirected hyperevents, where sender and receiver roles are not meaningful, one can set 8 and treat all participants as belonging to 9 (Lerner et al., 2019, Boschi et al., 8 Apr 2026).
The core rate specification is a Cox-type proportional-hazards model on hyperedges. For a candidate hyperedge 0 at time 1,
2
with relative rate
3
Equivalent formulations write
4
where 5 is a risk indicator, 6 a baseline hazard, and 7 an additive predictor built from hyperedge-specific covariates (Lerner et al., 2019, Boschi et al., 8 Apr 2026).
This generalization is substantively nontrivial. In multicast email, for example, a sender can repeatedly address the same receiver pair or larger receiver subset, and the suitability of one receiver can depend on which other receivers are included. In coauthorship, one publication is an event on a team of arbitrary size rather than a bundle of dyadic ties. In scientific publication networks, a single event can simultaneously involve a team of authors and a set of cited papers, and in tripartite extensions also a set of keywords. The model class therefore spans one-mode, two-mode, and multipartite higher-order event structures (Lerner et al., 2021, Lerner et al., 2023, Barbagli et al., 12 Apr 2026).
2. Risk sets, event choice, and likelihood structure
RHEMs are defined relative to a risk set 8: the set of hyperedges on which an event could occur at time 9. In temporal hypergraph formulations, if the observed event at index 0 has participant set 1, the candidate set is often conditioned on event size: 2 so the model explains which subset of that size occurs next, not how event size itself is generated. In publication-team models the same logic is used for coauthor teams; in mixed author-reference models the candidate set at time 3 becomes
4
so interpretation is conditional on the observed number of authors and references (Lerner et al., 2 Jun 2025, Lerner et al., 2021, Lerner et al., 2023).
The corresponding partial likelihood has the standard Cox event-choice form: 5 In the temporal-hypergraph formulation, the same logic appears as a sequential discrete-choice model: 6 with
7
This yields the familiar hazard-ratio interpretation: holding other statistics fixed, a one-unit increase in a hyperedge statistic multiplies the relative event rate by 8 (Lerner et al., 2019, Lerner et al., 2 Jun 2025).
A major modeling choice concerns whether risk sets are unconstrained or size-conditioned. In ministerial-meeting data, unconstrained models required explicit size and squared-size effects because event-size frequencies were highly nonuniform, whereas conditional-size models compared only participant substitutions within a fixed meeting size and yielded more stable interpretations of repetition and sub-repetition (Lerner et al., 2019). In coauthor and temporal-hypergraph work, conditioning on observed event size is presented as both substantively meaningful and computationally necessary, because the number of possible hyperedges grows explosively with node count and hyperedge size (Lerner et al., 2021, Lerner et al., 2 Jun 2025).
This conditioning is also the source of a common misconception. A positive coefficient in a size-conditioned RHEM does not mean that the process favors larger events per se. It means that, among candidate hyperedges of the same observed size, those with larger values of the statistic are more likely to be selected next (Lerner et al., 2 Jun 2025, Lerner et al., 2021).
3. Hyperedge statistics and higher-order dependence
The modeling language of RHEMs is a vector of temporal hyperedge statistics. These are candidate-event-level scores, not global summaries of the observed network. They quantify how a candidate hyperedge is positioned relative to the past.
A central example is subset repetition. For a temporal hypergraph with node set 9, prior history 0, and candidate hyperedge 1, prior degree of a node subset is
2
The subset-repetition statistic of order 3 is
4
For 5, this is the sum of prior node degrees and encodes preferential attachment of nodes; for 6, it favors candidate hyperedges containing dyads that have often co-occurred before; for 7, it favors repeated triples. Exact repetition is represented separately by
8
These constructions generalize preferential attachment from nodes to subsets of any order (Lerner et al., 2 Jun 2025).
In directed and sender-receiver-set settings, subset repetition must respect role structure. A canonical hyper-event analogue of repetition is
9
where 0 counts prior events in which all members of 1 jointly sent to all members of 2, possibly with additional participants. This is the hyper-event analogue of repetition, participation history, and local closure statistics in REMs (Boschi et al., 8 Apr 2026).
Beyond repetition, RHEMs include a broad class of higher-order dependence terms. In temporal hypergraphs, closure is defined by
3
This measures how much a candidate hyperedge would close prior two-paths. Homophily for a binary nodal attribute 4 is defined by
5
so a positive homophily coefficient implies that hyperedges with more similar members are favored. Degree assortativity of order 6 is defined by
7
with larger values corresponding to more similar degrees among the 8-subsets inside the candidate hyperedge (Lerner et al., 2 Jun 2025).
Polyadic communication models add genuinely hyperedge-level covariates that dyadic REMs cannot represent. These include receiver-set heterophily, exact repetition of the whole sender-plus-receiver-set configuration, unordered repetition of the participant set regardless of who is sender, partial receiver-set repetition for subset orders 9, sender-specific partial receiver-set repetition, and interaction among receivers. In the Enron email reanalysis, these statistics revealed stable conversational groups with turn-taking, repeated co-targeting of receiver pairs and triples, and sender-specific clustering of audiences—patterns that are unavailable to dyadic REMs because they concern dependence among receivers within the same event (Lerner et al., 2021).
Multipartite extensions redefine these statistics across node types. In tripartite publication models, closure terms are organized by endpoint types and intermediary type, producing within-set closures such as coauthoring closure or co-citation closure, and cross-part effects such as author-keyword-reference closure. The same work introduces Geometrically Weighted Subset Repetition (GWSR) to address steep growth, multicollinearity across orders, and lack of comparability across event sizes in standard subset repetition (Barbagli et al., 12 Apr 2026).
4. Estimation, sampled risk sets, and computational architecture
The main obstacle to RHEM estimation is not storing events but normalizing event rates over an enormous risk set. In dyadic REMs this denominator can already be prohibitive; in hyper-event models it grows combinatorially with the number of actors and allowable hyperedge sizes. For that reason, scalable estimation usually relies on sampled candidate sets rather than exhaustive evaluation (Lerner et al., 2019, Lerner et al., 2 Jun 2025).
A standard solution is case-control or nested case-control sampling. For each observed event hyperedge 0 or 1, one samples a fixed number of control hyperedges of the same size or same cardinality profile and replaces the full denominator by the sampled set. In temporal hypergraphs, for each observed hyperedge 2 one samples 3 alternatives of the same size uniformly at random from 4; in polyadic communication, one samples 5 alternative receiver sets of size 6 for the observed sender; in coauthor-team RHEMs, ten controls per event are sampled uniformly at random from the size-conditioned risk set (Lerner et al., 2 Jun 2025, Lerner et al., 2021, Lerner et al., 2021).
A crucial methodological point is that sampled likelihood need not imply approximate covariates. In scalable REM work that is directly relevant to RHEMs, all original events, not just sampled events, are still used to maintain the history 7. Sampling reduces the number of event times and controls entering the likelihood, but it does not induce approximate covariates when sufficient statistics are computed from the full event history. The same principle transfers directly to RHEMs: if focal hyper-events and sampled non-events are used only at estimation time, while history-dependent state variables continue to be updated from the full event log, the sufficient statistics remain correct (Lerner et al., 2019).
The statistical and computational consequences of this strategy are substantial. Dyadic REM work showed that a Cox-type REM can be fit on a network with 6.7 million users, 5.5 million articles, and 361,769,741 dyadic events by combining sampling of observed events with case-control sampling of non-events, and that tens of thousands of sampled events with a small number of controls per event can recover effect direction very reliably. For hyper-events, the same logic is even more compelling because the risk set is larger and the covariates are often sparser (Lerner et al., 2019).
Software implementations reflect this separation between state maintenance and estimation. eventnet is repeatedly used for efficient computation of explanatory variables for both REMs and RHEMs; coxph from R’s survival package is used for Cox partial likelihood and conditional logistic estimation; mgcv is used for generalized additive formulations with smooth and random effects (Boschi et al., 8 Apr 2026, Lerner et al., 2019, Boschi et al., 5 Sep 2025). This architecture—event-stream processing plus standard survival, conditional-logit, or GAM back ends—has become a practical template for large RHEM implementations.
5. Empirical domains and substantive findings
RHEMs have been applied across a wide range of higher-order relational domains. In government minister meetings, exact group repetition and lower-order familiarity mattered strongly, and event size played a dominant structural role. Conditional-size models for these data found positive exact repetition and positive sub-repetition of orders 1, 2, and 3, indicating that exact participant sets, familiar dyads, and familiar triads were all more likely to recur (Lerner et al., 2019).
In multicast email, RHEMs reveal dependencies among receivers that dyadic REMs miss. In the Enron reanalysis, many non-dyadic hyperedge effects were significant: unordered repetition was positive, exact repetition was negative conditional on unordered repetition, receiver-set heterophily was negative, higher-order subset repetition was positive, sender-specific higher-order subset repetition was positive, and interaction among receivers was positive. The full RHEM fit the data substantially better than the comparable dyadic model, with 8 versus 9. The joint interpretation of positive unordered repetition and negative exact repetition was stable conversational groups with turn-taking rather than repeated messages from the same sender to exactly the same set (Lerner et al., 2021).
Scientific collaboration has been a major testing ground because coauthorship is intrinsically polyadic. In one large application to three disciplines comprising 355,977 papers, RHEMs on coauthor teams showed positive 0 and 1, negative closure, negative 2, positive 3 and 4, and positive success disparity in the joint team-assembly model. The paired relational hyperevent outcome models showed that prior shared success increased both collaboration probability and impact, whereas some repetition and disparity effects diverged between team formation and paper performance (Lerner et al., 2021).
A coevolutionary RHEM for scientific publication networks treated each publication as one event simultaneously linking a set of authors to a set of cited papers. Applied to 1,416,353 papers and 1,286,941 unique authors, it found a positive tendency for subsets of papers to be repeatedly cited together across publications and identified “cite paper and its refs” as the strongest effect in the joint model. This result is methodologically important because it implies that papers’ citation impact may be partly due to endogenous network processes rather than only to paper-level attractiveness (Lerner et al., 2023).
Tripartite extensions carry the same logic further. In scientific collaboration networks modeled as events linking authors, references, and keywords, the full tripartite model had the best AIC, and omitting one node type changed not only coefficient magnitudes but in some cases signs and significance. The paper’s substantive claim was that publications should be treated as single higher-order events linking social participation, intellectual lineage, and semantic labeling, not as separate projections. It also introduced GWSR as a scalable repetition statistic for large multipartite event spaces (Barbagli et al., 12 Apr 2026).
RHEMs have also been used as tailored null models for temporal hypergraphs. In the Les Misérables chapter coappearance data, fitted temporal-hypergraph RHEMs showed preferential attachment of individual nodes, subset repetition of orders 2 and 3, triadic closure, assortativity of orders 1 and 2, and gender homophily. The paper’s main empirical lesson was that conclusions about higher-order overrepresentation depend on which lower-order structures are controlled for in the null model (Lerner et al., 2 Jun 2025).
6. Methodological issues, diagnostics, and recent extensions
Several methodological cautions recur across the literature. The first is that higher-order effects are easy to overinterpret. In dyadic REMs, unobserved sender and receiver heterogeneity can induce “ghost triadic effects,” especially transitive closure, even when no true triadic mechanism exists. The proposed remedy is a frailty or mixed-effects REM with sender and receiver random effects. This warning transfers directly to RHEMs: apparent subgroup closure, repeated team assembly, or higher-order participation effects may reflect latent actor differences in propensity to initiate, join, or attract hyper-events rather than genuine endogenous hyperedge mechanisms (Juozaitienė et al., 2022).
The second caution concerns sampled controls. Scalable REM work showed that rare statistics can be nearly degenerate on sampled non-events even when they appear well-behaved globally. In the Wikipedia study, repetition was nonzero for only about 6 in a million controls under uniform sampling, making its coefficient magnitude highly unstable. The general recommendation was to inspect the distribution of each statistic separately for events and sampled controls and, when necessary, replace naive uniform control sampling with stratified sampling or another design that enriches informative non-events. This issue is especially acute for RHEMs, where exact repeated group interactions and higher-order closure motifs are often even sparser (Lerner et al., 2019).
The third issue is model checking. For REMs with time-varying and random effects, a recent GOF framework uses weighted cumulative martingale residual processes and KS-type tests to assess whether covariates are correctly modeled, avoiding dependence on expensive event-sequence simulation. This approach was developed for dyadic REMs, but the underlying logic—compare observed statistics at event times to hazard-weighted expectations under the fitted model—is directly portable to RHEMs once hyperedge-indexed martingale residuals are available (Boschi et al., 2024).
Recent work has also relaxed the linearity and time-homogeneity assumptions of standard RHEMs. A generalized additive extension models hyperedge effects as smooth functions that can be nonlinear in the covariate and varying over calendar time, including tensor-product smooths: 5 Estimated through case-control partial likelihood recast as a GAM, this formulation detected threshold effects, saturation, reversals, and historically changing mechanisms in scientific collaboration and citation hyper-events that linear RHEMs would collapse into a single average slope (Boschi et al., 5 Sep 2025).
A final misconception concerns ontology. Multipartite publication events are often called “tripartite hyperevents,” but one paper explicitly notes that this is not standard hypergraph terminology because hypergraphs in formal graph theory typically live on a single node type. Methodologically, however, the same RHEM machinery—risk sets, hyperedge statistics, Cox-type partial likelihood, and sampled estimation—extends naturally to typed multipartite event structures once the admissible event space is defined (Barbagli et al., 12 Apr 2026).
Taken together, these developments position RHEMs as a broad Cox-style and discrete-choice-like modeling framework for dynamic hypergraphs and multipartite event systems. Their distinctive contribution is not merely to allow more than two participants, but to treat higher-order events as primitive stochastic objects, to define history dependence on sub-hyperedges and role-structured subsets, and to make inference about mechanisms that are invisible or distorted under dyadic projection (Lerner et al., 2019, Lerner et al., 2 Jun 2025).