---
title: 'MA-HERP: Hierarchical Motion-Action Prediction'
url: https://www.emergentmind.com/topics/ma-herp
type: topic
---

# MA-HERP: Hierarchical Motion-Action Prediction

MA-HERP, short for Movement-Action Hierarchical Estimation and Recursive Prediction, is a hierarchical and recursive probabilistic framework for the joint estimation and prediction of human movements and actions in human–robot collaboration [2604.03065]. It is designed for online settings in which a robot must reason simultaneously about continuous trajectories in $\mathbb{R}^d$ and discrete action labels such as reach, grasp, or release, rather than treating kinematics and symbolic intent as separate inference problems. Its defining elements are a hierarchy of action intervals linked by admissible Allen interval relations, a unified factorization coupling continuous dynamics, discrete labels, and explicit durations, and a Bayesian-filtering-inspired recursive inference scheme that alternates top-down action prediction with bottom-up sensory evidence [2604.03065].

## 1. Conceptual scope and problem formulation

MA-HERP addresses a characteristic difficulty of human–robot collaboration: safe and timely control depends on anticipating how the human body will move, while task-level coordination depends on inferring what the human intends to do next [2604.03065]. In this setting, movements supply continuous geometric information, whereas actions provide task semantics. The framework is explicitly motivated by the observation that realistic collaborative behaviour is not a single sequential stream: sub-actions can overlap across effectors, and composite actions impose temporal semantics such as “reach must occur before grasp” or “release may meet the end of a reach” [2604.03065].

The framework therefore represents actions as labeled time intervals rather than as isolated instantaneous symbols. An action interval is written
$$
a_k = (\ell_k, s_k, e_k),
$$
where $\ell_k \in L$ is a discrete label, $s_k$ and $e_k$ are integer start and end instants with $1 \le s_k < e_k \le T$, and the duration is
$$
d_k = e_k - s_k + 1.
$$
Continuous observations are denoted by $X_t \in \mathbb{R}^d$, optional context by $C_t$, and the induced instantaneous label at time $t$ by $\ell_t$ whenever $t \in [s_k,e_k]$ for some interval $a_k$ [2604.03065].

A central implication of this formulation is that intent recognition and motion forecasting become mutually informative rather than merely adjacent. Action-conditioned dynamics sharpen continuous forecasts, while the same dynamics yield data-dependent evidence for the discrete label posterior. The framework is therefore neither a purely symbolic temporal grammar nor a purely kinematic predictor; it is a coupled continuous–discrete model with explicit temporal structure.

## 2. Interval hierarchy and temporal admissibility

The temporal core of MA-HERP is Allen interval algebra. Given two intervals $a_i = (\ell_i, s_i, e_i)$ and $a_j = (\ell_j, s_j, e_j)$, their relation $R(a_i,a_j)$ belongs to the Allen set
$$
\mathcal{A} = \{\text{before}, \text{meets}, \text{overlaps}, \text{during}, \text{starts}, \text{finishes}, \text{equals}\},
$$
together with converses such as after, met-by, overlapped-by, contains, started-by, and finished-by [2604.03065]. The primitive relations are defined by the interval endpoints: before if $e_i < s_j$, meets if $e_i = s_j$, overlaps if $s_i < s_j < e_i < e_j$, during if $s_j < s_i$ and $e_i < e_j$, starts if $s_i = s_j$ and $e_i < e_j$, finishes if $s_i > s_j$ and $e_i = e_j$, and equals if $s_i = s_j$ and $e_i = e_j$ [2604.03065].

Not every relation is permitted for every pair of labels. MA-HERP encodes this through label-dependent admissibility sets
$$
\mathcal{A}_A(\ell_i,\ell_j) \subseteq \mathcal{A},
$$
which specify which relations are physically or task-wise plausible for the pair $(\ell_i,\ell_j)$ [2604.03065]. Hard constraints are expressed by plausibility factors
$$
\psi_R(a_i,a_j)=
\begin{cases}
1, & \text{if } R(a_i,a_j)\in \mathcal{A}_A(\ell_i,\ell_j),\\
0, & \text{otherwise,}
\end{cases}
$$
with the option for soft or fuzzy versions to tolerate segmentation noise [2604.03065]. The framework therefore supports partial order and controlled overlap rather than enforcing a strictly linear action chain.

Hierarchy is introduced through $H+1$ abstraction levels $h \in \{0,\dots,H\}$ [2604.03065]. Level $h=0$ contains movement segments aligned with continuous samples $X_{s_k^0:e_k^0}$, while levels $h \ge 1$ contain increasingly abstract actions. A level-$h$ action $a_k^h = (\ell_k^h, s_k^h, e_k^h)$ has a possibly empty child set $C(a_k^h) \subseteq A^{h-1}$, and a level-specific composition table
$$
M^h : L^{h-1} \times L^{h-1} \to 2^{\mathcal{A}}
$$
specifies which child-label pairs may stand in which relations [2604.03065]. A parent action is therefore interpreted as a temporally organized pattern of lower-level intervals, with consistency enforced by a composition factor $\psi_C(a_k^h, C(a_k^h))$.

This hierarchical representation is especially suitable for tasks in which semantic structure is carried by interval composition rather than by a flat state transition graph. The paper’s illustrative example is a handover whose children may require reach(object) before grasp(object), grasp(object) before reach(handover_pose), and reach(handover_pose) meets release(object) [2604.03065].

## 3. Unified probabilistic formulation

MA-HERP couples interval structure, explicit durations, and continuous dynamics in a single factorization. At the action level, it adopts a semi-Markov prior with explicit durations:
$$
p(\ell_{1:K}, s_{1:K}, e_{1:K} \mid C_{1:T})
=
\prod_{k=1}^K
p(\ell_k \mid \ell_{k-1}, C_{s_k}; \eta)\,
p(d_k \mid \ell_k; \eta)\,
\mathbf{1}[e_k = s_k + d_k - 1]\,
p(s_k \mid e_{k-1}),
$$
where $p(d_k \mid \ell_k)$ can be parameterized by a duration distribution or a hazard function $h_\ell(\cdot)$, as in HSMMs, to avoid memoryless durations [2604.03065].

Within each level, admissible Allen relations are imposed by products of $\psi_R$ over selected relevant interval pairs. Across levels, products of $\psi_C$ ensure that the children of each parent action satisfy the appropriate composition table $M^h$ [2604.03065]. The continuous component is action-conditioned:
$$
p(X_t \mid X_{t-1}, \ell_t, C_t; \theta),
$$
so that the instantaneous label $\ell_t$ selects the motion regime [2604.03065]. An extended state-space form is also allowed, with optional latent continuous state $h_t$, raw sensory measurement $y_t$, and regime variable $z_t$, although the reported experiments instantiate the model directly over $X_t$ without separate $y_t$ or $h_t$ [2604.03065].

Conditioned on context $C_{1:T}$, the full joint distribution factorizes as
$$
p(A^{(0:H)}, X_{1:T} \mid C_{1:T}) \propto
\left[\prod_{h=0}^H \prod_{(i,j)\in P^{(h)}} \psi_R(a_i^h, a_j^h)\right]
\left[\prod_{h=1}^H \prod_k \psi_C(a_k^h, C(a_k^h))\right]
\left[\prod_{t=1}^T p(X_t \mid X_{t-1}, \ell_t, C_t; \theta)\right]
\left[\prod_k p(\ell_k, d_k \mid \ell_{k-1}, C_{s_k:e_k}; \eta)\right].
$$
The paper states three corresponding conditional-independence assumptions: given $\ell_t$ and $X_{t-1}$, $X_t$ is independent of other labels and intervals; Allen factors couple only the start and end times of the affected interval pairs; and composition factors couple a parent interval with its children and the children among themselves via $M^h$ [2604.03065].

The continuous–discrete coupling operates in both directions. Top-down, the action-conditioned likelihood
$$
\lambda_t(\ell) \propto p(X_t \mid X_{t-1}, \ell, C_t; \theta)
$$
makes motion prediction label-specific. Bottom-up, the same likelihood acts as a data-dependent score for label inference, so the posterior over $\ell_t$ favours labels whose dynamics best explain the observed movement while satisfying duration and Allen constraints [2604.03065].

## 4. Recursive inference and computational profile

The online inference procedure is described as approximate message passing in filtering form rather than exact inference in the full factor graph [2604.03065]. For a single chain of labels, the framework defines $\pi_{t|t-1}(\ell)$ as the predictive prior, $\lambda_t(\ell)$ as the action-conditioned observation likelihood, and $\alpha_t(\ell)$ as the posterior belief over the instantaneous label. The recursion is
$$
\pi_{t|t-1}(\ell) \propto \sum_{\ell'} p(\ell \mid \ell', C_{t-1}; \eta)\,\alpha_{t-1}(\ell'),
$$
modulated in practice by the duration hazard $h_\ell(\cdot)$ to control boundary opening and closing,
$$
\lambda_t(\ell) \propto p(X_t \mid X_{t-1}, \ell, C_t; \theta),
$$
and
$$
\alpha_t(\ell) \propto \pi_{t|t-1}(\ell)\,\lambda_t(\ell)\,\Phi_{A,t}(\ell; \Pi_{t-1}),
$$
where $\Phi_{A,t}$ summarizes active Allen constraints and boundary or duration bookkeeping, and $\Pi_{t-1}$ denotes a summary of current interval hypotheses and associated statistics [2604.03065].

The per-step online algorithm is stated explicitly:

1. Predict the label prior $\pi_{t|t-1}(\ell)$ from $\alpha_{t-1}(\ell')$.
2. Score observations through $\lambda_t(\ell)$ from $p(X_t \mid X_{t-1}, \ell, C_t; \theta)$.
3. Enforce structure through $\alpha_t(\ell) \propto \pi_{t|t-1}(\ell)\lambda_t(\ell)\Phi_{A,t}(\ell; \Pi_{t-1})$.
4. Update durations and boundaries via the HSMM hazard $h_\ell$ and maintain $\Pi_t$.
5. Predict motion via
   $$
   \hat{X}_{t+1} = E[X_{t+1} \mid X_t, \ell_t^\star, C_{t+1}; \theta],
   $$
   with $\ell_t^\star$ the MAP label or a mixture if needed [2604.03065].

With multiple streams or hierarchy levels, $\ell$ becomes a tuple of labels and $\Phi_{A,t}$ factorizes over selected intra-stream and inter-stream pairs; composition constraints are updated when a parent hypothesis becomes sufficiently probable and its children’s partial ordering can be enforced [2604.03065].

The reported computational profile is
$$
O(|L|^2 + |L| \cdot \mathrm{cost}_{\mathrm{dyn}} + M^2)
$$
per time step for a single stream, where a dense transition matrix yields $O(|L|^2)$ for label prediction, likelihood evaluation costs $O(|L| \cdot \mathrm{cost}_{\mathrm{dyn}})$, Allen penalties over $M$ active intervals cost $O(M^2)$, and duration updates cost $O(|L|)$ [2604.03065]. Sparse transitions reduce the prediction term toward $O(|L|)$, and the paper notes that in online human–robot collaboration both $|L|$ and $M$ are typically modest.

## 5. Experimental instantiation and reported performance

The preliminary experimental evaluation instantiates MA-HERP on musculoskeletal simulations produced with Bioptim [2604.03065]. The simulated body is a torso and right arm with two shoulder DoFs and one elbow DoF, so $X_t \in \mathbb{R}^3$ consists of three joint angles. The task uses three target volumes $A$, $B$, and $C$, each a cube of $5$ cm side, together with a starting pose $I$. The label set contains nine transitions: $I \to A$, $I \to B$, $I \to C$, and all pairwise transitions among $A$, $B$, and $C$ [2604.03065]. For each label, the dataset contains $1000$ trajectories of length $T=251$ samples, with two noise regimes: D-0, which is noise-free, and D-0-10-30, which mixes $0\%$, $10\%$, and $30\%$ uniform additive noise along the trajectory [2604.03065].

The continuous component uses label-conditioned autoregressive Transformer predictors implementing $p(X_t \mid X_{t-1}, \ell, C_t; \theta)$, trained with history window $W=100$ and forecast horizon $P=50$, with inference by sliding window of stride $1$ [2604.03065]. D-0 uses $900$ train and $100$ test trajectories per label, while D-0-10-30 uses $2700$ train and $300$ test trajectories per label. The discrete component uses Transformer-based classifiers for the composite reach-to-target labels $I \to ABC$, $A \to BC$, $B \to AC$, and $C \to AB$; in the joint model, these classifiers provide the instantaneous belief $\alpha_t(\ell)$, and durations are known a priori in these experiments, so HSMM learning is not exercised [2604.03065].

For continuous motion prediction, the reported qualitative result is that D-0 forecasts are visually indistinguishable from ground truth, whereas D-0-10-30 yields more scattered forecasts under noise [2604.03065]. Quantitatively, D-0 achieves near-perfect PCC of approximately $0.98$–$0.99$ and RMSE of approximately $0.01$ rad for most labels, with occasional PCC drops on highly non-linear segments. Under D-0-10-30, PCC often falls to $0.3$–$0.8$ and RMSE rises to approximately $0.20$–$0.31$ rad, which the paper interprets as sensitivity to sensory perturbations and motivation for explicit uncertainty handling in human–robot collaboration [2604.03065].

Forecasting runtime for $P=50$ steps, measured on an Intel i7-10700 CPU at $2.90$ GHz, is reported as approximately $0.136$–$0.147$ s per forecast for D-0, with standard deviation $0.001$–$0.013$ s, and approximately $0.143$–$0.178$ s per forecast for D-0-10-30, with standard deviation $0.002$–$0.010$ s [2604.03065]. The paper states that these times are well below $1$ s and therefore compatible with online use.

For discrete action prediction, incremental prefix inference shows that $I \to ABC$ under D-0 saturates after approximately $170$ samples and under D-0-10-30 after approximately $230$ samples, while $A \to BC$ saturates after approximately $30$ samples in D-0 and approximately $130$ samples in D-0-10-30 [2604.03065]. $B \to AC$ and $C \to AB$ require longer evidence, and $B \to AC$ degrades more under noise. With a sliding window of length $10$ and majority vote over time, the reported accuracies are: for D-0, $I \to ABC$ $0.95$, $A \to BC$ $1.00$, $B \to AC$ $0.89$, and $C \to AB$ $0.61$; for D-0-10-30, $I \to ABC$ $0.66$, $A \to BC$ $1.00$, $B \to AC$ $0.50$, and $C \to AB$ $0.80$ [2604.03065]. The confusions are asymmetric: class $B$ recall can drop to approximately $0.15$ under noise despite high precision, making the classifier conservative in predicting $B$.

Classification runtime is approximately $0.6$–$0.7$ ms per prediction with negligible variance [2604.03065]. An end-to-end chaining experiment tracks a sequence of $15$ consecutive reaches in which the relations between consecutive reaches are before or meet; online beliefs $\alpha_t(\ell)$ update as evidence accrues, and the system produces continuous forecasts $\hat{X}$ and discrete labels consistent with the chain’s temporal semantics [2604.03065].

## 6. Relation to adjacent models, naming ambiguity, and open issues

MA-HERP is positioned against several neighboring modeling traditions [2604.03065]. Relative to separate pipelines consisting of a continuous predictor and an independent action classifier, it couples the two levels probabilistically: action-conditioned dynamics sharpen motion forecasts, motion likelihoods provide bottom-up evidence for actions, and Allen or duration constraints stabilize label inference and reduce implausible segmentations. Relative to HMM or HSMM formulations, it preserves HSMM-like durations but augments them with Allen interval constraints capable of representing meets, overlaps, starts, and finishes rather than only sequential transitions. Relative to Switching Linear Dynamical Systems, it replaces linear-Gaussian mode dynamics with non-linear label-conditioned neural dynamics and adds hierarchical interval structure with explicit Allen relations and composition tables [2604.03065].

The framework rests on several stated assumptions. It assumes a fixed or slowly varying set of admissible Allen relations $\mathcal{A}_A(\ell_i,\ell_j)$, stationary semi-Markov durations or hazards within a task context, time-aligned observations $X_t$, and available optional context $C_t$ [2604.03065]. The paper also identifies several limitations: the current experiments use a simplified $3$-DoF reaching setup with a limited action vocabulary, do not provide end-to-end closed-loop human–robot collaboration validation under realistic sensing latency, occlusions, or multimodal fusion, leave uncertainty calibration of the continuous predictors to future work, and note that the interaction between history length $W$ and horizon $P$ requires systematic analysis to avoid oscillations and over-commitment [2604.03065].

The proposed extensions are correspondingly directed toward richer temporal, sensory, and interaction structure. They include multi-agent settings with cross-agent Allen relations and joint composition tables, multimodal sensing through a richer observation model $p(y_t \mid X_t)$ and learned sensor fusion, richer action grammars and symbolic task constraints through expanded composition tables and logic-based priors, and explicit latent state $h_t$ with probabilistic emission models for uncertainty propagation [2604.03065]. A plausible implication is that MA-HERP is intended less as a fixed architecture than as a template for coupled continuous–discrete online inference in temporally structured collaboration tasks.

A separate terminological issue arises because the label “MA-HERP” is not used uniformly across the broader literature represented here. In the proteomics paper “HERP: Hardware for Energy Efficient and Realtime DB Search and Cluster Expansion in Proteomics,” the term “MA-HERP” is not explicitly defined; the text instead presents “Memory-Augmented HERP” as a reasonable interpretation for a proposed extension of HERP’s compute-in-memory architecture [2511.03437]. By contrast, the explicit and fully specified definition of MA-HERP as Movement-Action Hierarchical Estimation and Recursive Prediction belongs to the human–robot collaboration framework summarized above [2604.03065].

Source: https://www.emergentmind.com/topics/ma-herp