Papers
Topics
Authors
Recent
Search
2000 character limit reached

ChimpACT: Longitudinal Chimpanzee Behavior Dataset

Updated 15 July 2026
  • ChimpACT is a longitudinal video dataset comprising 163 clips from Leipzig Zoo, Germany, enabling in-depth analysis of chimpanzee behavior and social dynamics.
  • The dataset provides richly annotated frames with detection, identification, 2D pose keypoints, and 23 fine-grained behavior classes over 23 individuals.
  • It underpins multi-task benchmarks for tracking, pose estimation, and action detection and supports domain-adaptive pretraining and zero-shot tracking studies.

ChimpACT is a longitudinal video dataset for understanding chimpanzee behaviors in a stable social group at the Leipzig Zoo, Germany. It was introduced to address the lack of datasets on non-human primate behavior and to support detection, identification, pose estimation, and fine-grained spatiotemporal behavior analysis in group-living chimpanzees. The dataset spans from 2015 to 2018, centers on a social group of 23 chimpanzees, and includes a particular focus on the developmental trajectory of one young male, Azibo. Across 163 videos and 160,500 frames, ChimpACT provides richly annotated detection, identification, pose, and behavior labels, and has subsequently been used both as a benchmark and as in-domain data for domain-adaptive pretraining and zero-shot multi-animal tracking studies (Ma et al., 2023).

1. Dataset scope and longitudinal design

ChimpACT was collected at the Wolfgang Köhler Primate Research Center, Leipzig Zoo, Germany, from 2015 to 2018 during daytime hours from 7 am to 4 pm. The subjects comprise a stable social group of 23 chimpanzees (Pan troglodytes), including mothers, juveniles, and adults. A defining feature is longitudinal “focal sampling” of one male infant, Azibo, who was tracked from birth through age 3. The recordings were made in semi-naturalistic indoor and outdoor enclosures measuring 400 m² and 4,000 m², respectively, with climbing structures, hammocks, vegetation, foraging boxes, and an artificial river (Ma et al., 2023).

The dataset contains 163 clips with a cumulative 160,500 frames, corresponding to approximately 2 hours of video. The dataset paper reports a frame rate of 25 fps with H.264 encoding, two image resolutions—720×578 and 1,280×720 pixels—and clip lengths of approximately 1,000 frames each. A tripod-mounted JVC Everio camera with optical zoom and panning was used. The same source states that pose and action labels are annotated every 10th frame, while detection and tracking are annotated on all frames (Ma et al., 2023).

The longitudinal design uniquely captures ontogeny of social skills and cultural behaviors in Azibo, formation of dominance and kinship relationships, and changes in social interaction patterns as Azibo matures. This makes ChimpACT relevant not only for benchmark construction in computer vision but also for comparative social cognition, welfare analysis, and longitudinal studies of group structure. A plausible implication is that the persistent recording of a stable social group enables analyses that are difficult to recover from short-duration or cross-sectional primate datasets.

2. Annotation schema and ethogram

For every annotated frame, ChimpACT provides a detection bounding box with a visibility flag, a tracking ID consistent within a clip, a real identity label drawn from 23 unique names, 2D body-pose keypoints, and spatiotemporal behavior labels. Visibility is labeled as fully visible, truncated, or occluded. The 2D pose annotation uses 16 joints with keypoint visibility following COCO conventions, where visibility takes values $0$ for outside frame, $1$ for inside but occluded, and $2$ for clearly visible (Ma et al., 2023).

The 16 keypoints are defined as follows: $0$ root of hip; $1$ right knee; $2$ right ankle; $3$ left knee; $4$ left ankle; $5$ neck; $6$ upper lip; $1$0 lower lip; $1$1 right eye; $1$2 left eye; $1$3 right shoulder; $1$4 right elbow; $1$5 right wrist; $1$6 left shoulder; $1$7 left elbow; and $1$8 left wrist. Identity labels for 23 individuals were confirmed by a senior primatologist, and the identity distribution is long-tailed, with Azibo and his mother Swela the most frequent (Ma et al., 2023).

The behavior annotation defines 23 fine-grained classes organized into four top-level categories. These are locomotion, object interaction, social interaction, and others. The locomotion category contains four subcategories: moving, climbing, resting, and sleeping. Object interaction contains four subcategories including solitary playing, eating, and manipulating. Social interaction contains 12 subcategories, including grooming/being groomed, aggression, embracing, begging/being begged from, taking/losing objects, carrying/being carried, nursing/being nursed, playing, and touching. The others category contains erection and displaying. Multiple labels per individual per frame are allowed, and role distinctions such as “grooming” versus “being groomed” are explicitly annotated (Ma et al., 2023).

The annotation workflow used BasicFinder CO., Ltd.’s private labeling tool, 15 trained annotators, and 2 project managers. The process included comprehensive documentation with visual exemplars, a pilot phase with trial annotations and feedback, expert review by in-house primatologists for all identity and behavior labels, and iterative quality control in which labels failing audits were returned for correction. The total annotation cost is reported as approximately 70,000 RMB (Ma et al., 2023).

3. Benchmark tasks and evaluation protocols

ChimpACT defines three benchmark tracks: multi-object tracking and re-identification, 2D pose estimation, and spatiotemporal action detection. The dataset paper evaluates methods on an 80/10/10 train/val/test split with all identities present in each split for tracking, and standard train/val/test splits for pose estimation and action detection (Ma et al., 2023).

Track Representative methods Representative result
Multi-object tracking & re-identification Faster R–CNN, YOLOX, SORT, DeepSORT, Tracktor, QDTrack, ByteTrack, OC-SORT OC-SORT + YOLOX: HOTA ≈ 47.9%, MOTA ≈ 42.1%, IDF1 ≈ 53.3%, mAP ≈ 70.5%
2D pose estimation CPM, SimpleBaseline, RLE, StackedHourglass, MobileNetV2, HRNet-W32/W48, DarkPose, HRFormer DarkPose w/ HRNet-W32: [email protected] ≈ 65.6%, AP ≈ 25.9%, AP$1$9 ≈ 58.2%
Spatiotemporal action detection ACRN, LFB, SlowOnly, SlowFast SlowFast: overall mAP ≈ 24.3%, $2$0, $2$1, $2$2, $2$3

For tracking, the reported metrics include mean Average Precision, MOTA, IDF1, HOTA, and normalized FP, FN, and ID switches. The dataset paper gives the formulas

$2$4

and

$2$5

For pose estimation, the principal metrics are PCK@$2$6 and AP over keypoint heatmaps at thresholds such as 0.5 and 0.75, with

$2$7

For action detection, the central metric is mAP over 23 action classes, with breakdown by top-level categories $2$8, $2$9, $0$0, and $0$1 (Ma et al., 2023).

These benchmarks make clear that ChimpACT is designed as a multi-task resource rather than a single-purpose action dataset. This suggests that methodological advances in one track, such as improved tracking or representation learning, can plausibly propagate to other tracks through shared visual structure and shared annotation coverage.

4. Domain-adaptive pretraining for action recognition

A later study uses ChimpACT as in-domain data for self-supervised domain-adaptive pretraining (DAP) of a V-JEPA model and reports improved action recognition on primate behavior. In that account, ChimpACT is described as 2 hours of zoo-housed chimpanzee footage covering approximately 20 individuals, with 23 distinct behavior classes annotated on a per-frame, multi-label basis and divided into indoor and outdoor sessions. The same source groups the 23 classes conceptually into locomotion, object interaction, social interaction, and other (Mueller et al., 15 Sep 2025).

For DAP, the raw videos are split into overlapping 3 s snippets with stride $0$2 s, yielding approximately 6,000 samples. Each snippet $0$3 is processed by a zero-shot Grounding DINO detector with text prompt “monkey.primate.ape.” and threshold $0$4 to localize one or more chimpanzee bounding boxes; snippets with no detection are discarded. A randomly selected box $0$5 is enlarged by 25% and spatially jittered with scale 0.3–1.0 and aspect 0.75–1.35. The pipeline then applies $0$6 spatial frames and $0$7 temporally subsampled frames with stride $0$8, yielding a tensor $0$9 (Mueller et al., 15 Sep 2025).

The backbone is V-JEPA, with encoder $1$0 implemented as a ViT-L with 304 M parameters, predictor $1$1 as a 12-layer transformer with 22 M parameters, and target encoder $1$2 as an EMA of $1$3. The input $1$4 is patchified into $1$5 tokens $1$6 with $1$7. The self-supervised objective is the masked latent reconstruction loss

$1$8

DAP initializes $1$9, $2$0, and $2$1 from publicly released V-JEPA weights pretrained on human video from HowTo100M, Kinetics, and SSv2, and then continues minimizing $2$2 on in-domain ChimpACT snippets without labels. Reported hyperparameters are 14,400 gradient steps, effective mini-batch size 80, approximately 1.2 M samples seen, learning rate $2$3 with linear warmup and cosine decay, and weight decay ramped from 0.01 to 0.1. Training requires approximately 3 h on 4 × A100 GPUs (Mueller et al., 15 Sep 2025).

For downstream evaluation on ChimpACT, the encoder $2$4 is frozen. For each frame-level ground-truth box $2$5 at time $2$6, a 2 s snippet centered at $2$7 is cropped and passed through $2$8 and $2$9, producing $3$0. The classifier head uses one multi-head cross-attention from a learnable query $3$1 to keys and values $3$2, followed by a linear projection to $3$3 logits $3$4. Because behavior tags are multi-label, optimization uses binary cross-entropy,

$3$5

Training runs for up to 30 epochs with early stopping on validation mAP, and only the classifier’s 4 M parameters are learned (Mueller et al., 15 Sep 2025).

On the ChimpACT test set, the reported mAP values are 24.4 for ACRN, 24.3 for SlowFast, 26.0 for V-JEPA without DAP, and 30.7 for V-JEPA + DAP. The per-group results are $3$6; $3$7; and $3$8, respectively. Compared to the off-the-shelf V-JEPA, DAP yields $3$9 mAP overall, with object-interaction mAP increasing by $4$0 points and social-interaction mAP by $4$1 points, while locomotion drops by $4$2 points (Mueller et al., 15 Sep 2025).

The same study states that behavior classes with sparse representation, including social play and tool use, benefit disproportionately from DAP, and that a classifier-head ablation showed a single cross-attention layer with 4 M parameters matches or outperforms the original 12 M-parameter block. It also identifies limitations: per-box cropping discards global context and incurs multiple forward passes for multi-instance frames, and the drop in $4$3 indicates that DAP must be balanced if locomotion detection is critical. The interpretation offered there is that generic pretraining already captures gross motion well while DAP sharpens fine-grained manipulative and social cues (Mueller et al., 15 Sep 2025).

5. Zero-shot multi-animal tracking on ChimpACT

ChimpACT has also been used to evaluate zero-shot multi-animal tracking. In that setting, the dataset is described as 163 video sequences of free-moving chimpanzees recorded at 25 FPS in the Leipzig Zoo, with the test set alone comprising roughly 160.8 K frames, 712 individual track identities, and 563 K annotated bounding boxes. Each chimpanzee is annotated with a tight axis-aligned bounding box and a unique track ID across the sequence. Reported challenges include variable lighting and background due to outdoor enclosures, substantial variation in chimpanzee pose, frequent partial and full occlusion when individuals group together, and low object scale relative to image resolution of 576p–720p (Meier et al., 4 Nov 2025).

The zero-shot tracker combines a pretrained Grounding DINO (Swin-L) detector with the Segment Anything Model 2 (SAM 2) tracker, plus three heuristics—adaptive thresholding, mask-based initialization, and density-aware reconstruction—and mask NMS. All hyperparameters, including text prompt “ape” for Grounding DINO, are held constant, and no dataset-specific fine-tuning is performed (Meier et al., 4 Nov 2025).

Adaptive thresholding begins with Grounding DINO confidence scores $4$4, partitions the 1D score distribution into two clusters via K-Means with $4$5,

$4$6

and sets a per-sequence threshold

$4$7

with static offset $4$8. Detections with $4$9 pass to tracking. For track association, each retained detection box $5$0 is matched to existing tracks by bipartite matching on box IoU,

$5$1

Unmatched detections are refined through SAM 2 segmentation: the normalized mask intersection is computed as

$5$2

and a new track is initialized only if $5$3 with $5$4 (Meier et al., 4 Nov 2025).

Density-aware reconstruction periodically re-prompts a track with its most reliable detection box only when the match is unambiguous. For a candidate box $5$5, letting $5$6, the criterion is

$5$7

with $5$8. Track termination uses SAM 2 occlusion scores $5$9: tracks with $6$0 for $6$1 consecutive frames are terminated. Duplicate masks are suppressed via mask-level NMS with $6$2 (Meier et al., 4 Nov 2025).

On the ChimpACT test set, the reported performance is HOTA 58.6, DetA 49.8, AssA 70.1, DetRe 57.3, LocA 83.4, MOTA 48.6, IDF1 66.7, and IDSW 32. Relative to trained AlphaChimp, which attains HOTA 56.3 with associations unreported, the zero-shot method improves HOTA by $6$3 points without any fine-tuning or hyperparameter adjustment. Relative to zero-shot baselines using the same Grounding DINO inputs, it improves over ByteTrack by $6$4 HOTA and $6$5 AssA, and over NetTrack by $6$6 HOTA and $6$7 AssA (Meier et al., 4 Nov 2025).

The reported strengths are high association accuracy, reduced false positives and premature track splits from adaptive thresholding and mask-based initialization, and zero-shot generalization across four datasets. Reported limitations are modest recall, linear growth of SAM 2 runtime and GPU memory consumption with the number of tracks, and residual ID switches during prolonged full-body occlusion or very rapid out-of-plane motion. Recommendations for future work include integrating a lightweight motion model, incorporating semantic cues or appearance embeddings, and exploring GPU–CPU offloading and model pruning (Meier et al., 4 Nov 2025).

6. Technical challenges, scientific significance, and open directions

The dataset paper identifies several technical challenges: low color-contrast fur, frequent self-occlusions in group settings, long-tail distributions for rare behaviors and individuals, complex 3D articulations not well represented by human-centric 2D models, and fine distinctions in social actions such as “touching” versus “grooming” (Ma et al., 2023). Later tracking work reiterates related difficulties through outdoor lighting variation, low object scale, and frequent occlusion, while later representation-learning work highlights the sensitivity of action performance to domain mismatch and to the preservation or loss of global context (Meier et al., 4 Nov 2025).

ChimpACT’s scientific significance is framed in several ways. It bridges the gap between human-centric vision research and non-human primate ethology, enables quantitative studies of social network formation, cultural transmission, and welfare, and supports downstream analyses such as modeling dominance hierarchies via longitudinal interaction graphs, computational ethology through unsupervised discovery of new behavioral motifs, and cross-species transfer learning to macaques, marmosets, or human infant studies (Ma et al., 2023). The longitudinal focus on Azibo also provides direct access to developmental changes, including transition from mother-dependence through nursing and carrying to peer play, while long-term tracking highlights the emergence of individual movement patterns and social bonds (Ma et al., 2023).

Common misconceptions can be addressed directly from the benchmark evidence. ChimpACT is not only an action-recognition dataset; it is explicitly structured around tracking and re-identification, 2D pose estimation, and spatiotemporal action detection (Ma et al., 2023). It is also not restricted to supervised learning use cases: later studies use it for label-free domain-adaptive pretraining and for zero-shot tracking without dataset-specific fine-tuning (Mueller et al., 15 Sep 2025). Conversely, the benchmark results do not support the view that the dataset has become solved. Reported performance remains limited in several respects, including modest action-detection mAP, recall constraints in zero-shot tracking, and persistent sensitivity to rare classes, occlusion, and global context (Mueller et al., 15 Sep 2025).

Future directions stated in the sources include pose-tracking and 3D reconstruction via multi-view or depth sensors, few-shot or zero-shot learning for rare behaviors in the “others” category, weakly or semi-supervised methods to reduce annotation effort, multimodal fusion with audio or RFID, scaling DAP across multiple animal domains, and exploring joint finetuning of encoder and head to recover lost global-motion sensitivity (Ma et al., 2023). Taken together, these directions indicate that ChimpACT functions both as a benchmark suite and as a substrate for broader research on representation learning, computational ethology, and longitudinal analysis of chimpanzee sociality.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (3)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ChimpACT.