Papers
Topics
Authors
Recent
Search
2000 character limit reached

Pyramidal Shapley-Taylor Framework

Updated 5 February 2026
  • Pyramidal Shapley-Taylor Learning Framework is a method that enables fine-grained motion-language retrieval through hierarchical token compression and pyramidal alignment.
  • It leverages multi-stage contrastive learning and joint-segment alignment to improve accuracy in matching detailed human motion with natural language.
  • The framework incorporates Shapley-Taylor interaction attribution, providing enhanced interpretability by quantifying specific contributions of motion and text tokens.

The Pyramidal Shapley-Taylor (PST) Learning Framework is a methodology for fine-grained motion-language retrieval that models hierarchical cross-modal alignment between human motion sequences and natural language. It departs from global-centric paradigms by adopting a pyramidal process reflecting human motion perception, leveraging Shapley-Taylor interaction attribution for interpretability and precise alignment. PST combines hierarchical token compression, contrastive learning, and Shapley-Taylor-based interaction distillation, and demonstrates superior retrieval accuracy and interpretability on standard benchmarks (Chen et al., 29 Jan 2026).

1. Hierarchical Motion and Text Decomposition

PST represents a human motion sequence as a set of frames, each comprising JJ body joints in CC-dimensional coordinates (C=3C=3 for 3D skeletons). Let M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L denote a sequence of LL frames. At the base, the motion is flattened into Nint=J⋅LN_{int} = J \cdot L joint-stage tokens J={jk}k=1...Nint\mathcal{J} = \{ j_k \}_{k=1...N_{int}}. These tokens are compressed into NsgmN_{sgm} segment tokens S={si}i=1...Nsgm\mathcal{S} = \{ s_i \}_{i=1...N_{sgm}} using a token compressor that integrates convolutional layers, self-attention, and KNN-DPC clustering, with compression ratio p=Nsgm/Nintp = N_{sgm} / N_{int} (typically CC0). Segments are pooled further to generate a single global descriptor CC1 for the motion.

The textual description CC2 is tokenized into word tokens CC3, compressed via an analogous token compressor into phrase-stage tokens CC4, and then pooled into a global text feature CC5.

This multi-stage tokenization underpins the pyramidal processing of representation granularity: from joint/word-level, to segment/phrase-level, to global motion/text features.

2. Shapley-Taylor Interaction Attribution

Central to PST is the quantification of fine-grained cross-modal interactions using Shapley-Taylor interaction (STI) [Sundararajan et al., 2020]. For each pair of motion and text tokens CC6, the second-order (CC7) Shapley-Taylor interaction index CC8 is computed as the permutation-average contribution of including tokens CC9 (from motion) and C=3C=30 (from text) to the final retrieval scoring function C=3C=31. Mathematically, for C=3C=32 tokens in total and for each permutation C=3C=33, define C=3C=34 as the set preceding both C=3C=35 and C=3C=36; then,

C=3C=37

Direct computation is intractable; thus, PST introduces an STI Estimation Head C=3C=38 that approximates C=3C=39 via Monte-Carlo sampling and is trained by minimizing

M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L0

where M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L1 and M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L2 denote the softmax distributions of true and predicted M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L3 values, respectively, facilitating efficient end-to-end learning of interaction attributions between fine-level tokens.

3. Pyramidal Multi-Level Alignment Strategy

PST's core mechanism is structured as a three-stage pyramid, each supervised with local contrastive objectives and interaction distillation:

  • Joint-Wise Alignment: At the lowest pyramid level, joint-stage motion tokens and word tokens are compared. Pairwise cosine similarities M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L4 are computed after projection, and an InfoNCE contrastive loss M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L5 is applied. STI distillation (M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L6) further aligns the STI Head with true Shapley-Taylor indices at this local scale.
  • Segment-Wise Alignment: Tokens are compressed into segments/phrases. Segment-stage similarities M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L7 are calculated; InfoNCE loss M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L8 and STI distillation M={mt∈RJ×C}t=1LM = \{ m_t \in \mathbb{R}^{J \times C} \}_{t=1}^L9 are used. A consistency loss

LL0

enforces knowledge distillation between levels.

  • Holistic Alignment: At the top, segment tokens are pooled to global motion/text descriptors; global similarity LL1 is used in LL2. STI distillation is not performed at this level.

At each level, token transformation and compression leverage the token compressor stack—convolutional layers, LayerNorm, multi-head self-attention, KNN-DPC clustering, and further self-attention—to efficiently represent increasing abstraction.

4. Objective Function and Optimization

The total loss for PST integrates multiple objectives: LL3 where LL4, LL5, LL6 control the pyramid level weights, LL7 weight the STI losses, and LL8 the knowledge distillation. Each contrastive loss LL9 adopts the InfoNCE formulation, measuring cross-modal retrieval accuracy within a batch: Nint=J⋅LN_{int} = J \cdot L0 where Nint=J⋅LN_{int} = J \cdot L1 is the temperature.

5. Model Architecture Components

PST comprises distinct architectural modules:

Module Architecture Description
Motion Encoder Vision Transformer (ViT) on spatio-temporal MotionPatches
Text Encoder DistilBERT for contextual word embeddings
Projection Head Two-layer MLP (GeLU activation) to unify token feature space
Token Compressor Conv(3×1) → LayerNorm → Self-Attention → KNN-DPC clustering → (repeat)
STI Estimation Head Conv(3×3) → ReLU → Self-Attention → Residual → Conv(3×3) → ReLU

The repeated application of the token compressor enables hierarchical abstraction. The STI Head is trained to predict Shapley-Taylor indices for pairs of tokens.

6. Empirical Performance

On standard motion-language retrieval benchmarks, PST demonstrates superior performance compared to prior state-of-the-art approaches using only global alignment or limited part awareness. On the HumanML3D dataset under the "All" protocol, Text→Motion Recall@1 improves from 10.80 to 12.45, and Motion→Text Recall@1 from 71.61 to 76.15 compared to MotionPatch. On KIT-ML, PST achieves Recall@1 of 16.01 (Text→Motion) versus 14.02 and Recall@1 of 56.83 (Motion→Text) versus 53.55. These improvements are consistent across various batch protocols and across broader metrics including Recall@2, @3, @5, @10, and MedR (Chen et al., 29 Jan 2026).

7. Interpretability via Shapley-Taylor Attribution

PST's explicit computation of Shapley-Taylor indices Nint=J⋅LN_{int} = J \cdot L2 provides intrinsic interpretability. High Nint=J⋅LN_{int} = J \cdot L3 values indicate strong correspondence between specific joint movements and linguistic tokens (e.g., "right knee bends" and "kneels"), both at single-joint/word and segment/phrase scales. Segment-level attributions reveal mapping between joint clusters and multi-word expressions. Visual heatmaps of Nint=J⋅LN_{int} = J \cdot L4 scores elucidate temporal and spatial loci of model attention, enabling fine-grained diagnostic analyses and insights into the model's reasoning. This attribution-based transparency is absent from purely global contrastive frameworks, positioning PST as both empirically strong and interpretable (Chen et al., 29 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Pyramidal Shapley-Taylor (PST) Learning Framework.