Papers
Topics
Authors
Recent
Search
2000 character limit reached

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

Published 12 Aug 2026 in cs.CV | (2608.11574v1)

Abstract: Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these applications, the visibility of each hand joint in the image is crucial for assessing the reliability of estimation results under occlusion. However, most existing HPE methods output joint positions without explicitly indicating their visibility. Although some methods account for occlusion or visibility, visibility estimation has mainly been used as an auxiliary signal for improving pose estimation. To our knowledge, per-joint hand visibility estimation has not been systematically studied as a standalone task. In this work, we propose Hand Visibility Detector, a model for estimating the visibility of individual hand joints, and present the first systematic investigation of visibility estimation as an independent task. We show that leveraging the prior knowledge of HPE models pretrained on large-scale data as a backbone yields high performance in this task. We further demonstrate the utility of Hand Visibility Detector on a downstream task of 3D hand pose annotation via multi-view triangulation of 2D keypoints, showing that visibility-weighted triangulation reduces reprojection error. Our method is released as a ready-to-use package, and the code and demo are available at https://github.com/ryhara/hand_visibility_detector .

Summary

  • The paper introduces a lightweight GAU-based detector that predicts visibility for all 21 MANO keypoints using frozen HaMeR or WiLoR hand-pose features.
  • The method achieves 0.931 mAP and 0.896 F1 on HInt, outperforming ImageNet-pretrained baselines by 3.4 mAP points with only 0.83 million trainable parameters.
  • Visibility-weighted multi-view triangulation reduces reprojection errors across DexYCB, HO3D, and H2O, with a maximum mean-error reduction of 10.1% on HO3D.

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

Problem Formulation and Motivation

โ€œHand Visibility Detector: Per-Keypoint Visibility Estimation for Handsโ€ (2608.11574) formulates per-keypoint hand visibility estimation as an independent computer-vision task rather than as an auxiliary component of hand pose estimation. The distinction is technically important. Existing HPE systems generally produce 2D or 3D joint locations even when individual joints are occluded, truncated by the image boundary, or visually ambiguous. Their outputs therefore conflate two different quantities: the estimated position of a joint and the degree to which that joint is directly observable.

The paper defines visibility for each of the 21 MANO keypoints as a binary property: a joint is visible if it is neither occluded nor outside the image frame. The proposed model predicts a continuous confidence score in [0,1][0,1] for every joint. This formulation is intended to provide an explicit reliability signal for downstream systems, including multi-view 3D annotation, robotic perception, AR/VR interaction, and uncertainty-aware pose tracking.

The central empirical claim is that large-scale HPE pretraining provides a substantially more suitable representation for visibility estimation than generic ImageNet pretraining, even though the visibility head itself is lightweight and trained on a comparatively modest manually annotated dataset.

Figure 1

Figure 1: Per-joint visibility scores distinguish directly observable hand joints from joints occluded or truncated in the image.

Architecture

The model receives a cropped hand image and consists of two components: a frozen ViT-based hand encoder and a trainable visibility head. The encoder is initialized from either HaMeR or WiLoR, both large-scale hand reconstruction systems trained to infer 3D hand structure under diverse imaging conditions. Its output tokens are rearranged into a spatial feature map of dimensions 16ร—12ร—128016 \times 12 \times 1280.

Freezing the encoder is a deliberate design choice. Visibility estimation requires reasoning about hand topology, inter-joint relationships, occlusion patterns, and image context. These capabilities are assumed to be encoded in the pretrained HPE representation. Training only the task-specific head also substantially reduces optimization cost and limits overfitting to the visibility annotations in HInt.

The visibility head follows the general structure of the RTMPose head. A 1ร—11 \times 1 convolution compresses the feature dimensionality, after which the spatial tokens are processed by a Gated Attention Unit (GAU). The GAU is intended to model global dependencies across the image rather than make independent local decisions for each spatial region. A final projection produces 21 joint-specific logits, which are spatially averaged and transformed with sigmoid functions into visibility probabilities. Training uses binary cross-entropy over the 21 joints.

Figure 2

Figure 2: The pipeline combines a frozen pretrained HPE ViT backbone with a lightweight GAU-based per-joint visibility head.

The complete model contains 631 million parameters, but only 0.83 million parameters are trained, corresponding to 0.131% of the full network. Training takes approximately 2.5 hours on a single NVIDIA H200 GPU with roughly 10 GB of memory. This parameter-efficient configuration is relevant for practical deployment because it permits the addition of visibility prediction without modifying or retraining the underlying HPE system.

Dataset and Evaluation Protocol

Training and evaluation use HInt, a dataset containing manually annotated per-joint visibility labels across diverse in-the-wild imagery, including web images and egocentric video frames. The experimental split contains 25,273 training frames and 5,374 evaluation frames. The dataset is particularly important to the paperโ€™s argument because earlier visibility-aware hand models were generally trained in controlled settings and were not evaluated as standalone visibility predictors under heterogeneous occlusion conditions.

The authors report mean average precision (mAP) and F1 score. mAP evaluates ranking quality over continuous visibility scores, while F1 evaluates binary predictions after thresholding at 0.5. The comparison includes the visibility estimator from Kim et al. and the estimator used in Contact4D, both of which rely on ImageNet-pretrained CNN representations rather than hand-specific HPE backbones.

The proposed model obtains an mAP of 0.931 and an F1 score of 0.896. The Kim et al. baseline reaches 0.895 mAP and 0.858 F1, while the Contact4D baseline reaches 0.897 mAP and 0.860 F1. Thus, the improvement is 3.4 mAP points over the strongest baseline, matching the paperโ€™s principal quantitative claim.

Representation Transfer and Ablation Results

The backbone ablation provides the strongest evidence for the paperโ€™s representation-learning hypothesis. Generic CNN backbones perform substantially worse: CSPNeXt-X obtains 0.800 mAP and 0.799 F1, while ResNet-152 obtains 0.796 mAP and 0.797 F1. A generic ViT-H improves performance to 0.838 mAP and 0.819 F1, and DINOv3 reaches 0.897 mAP and 0.866 F1. However, the hand-specific HaMeR and WiLoR encoders achieve 0.932 and 0.931 mAP, respectively, with both reaching approximately 0.896 F1.

This result indicates that model scale alone is insufficient. The generic DINOv3 representation is substantially larger than the hand-specific encoders, yet remains inferior. The relevant prior is not merely visual feature quality but structured knowledge of hand geometry, articulation, and occlusion acquired through HPE training.

The paper also reports an apparently counterintuitive result: fine-tuning the WiLoR backbone for visibility estimation reduces mAP from 0.931 to 0.622 and F1 from 0.896 to 0.704. The authors interpret this as evidence that task-specific fine-tuning corrupts useful pretrained representations. The claim is consequential but should be interpreted narrowly. The experiment fine-tunes under the stated training protocol and dataset scale; it does not establish that all forms of fine-tuning are harmful. More conservative adaptation strategies, such as low-rank adaptation, discriminative learning rates, partial unfreezing, stronger regularization, or multi-task objectives, could potentially avoid this degradation.

The visibility-head ablation further supports the architecture. Removing the GAU lowers mAP from 0.931 to 0.887 and F1 from 0.896 to 0.860. Replacing the proposed head with the linear head used by Kim et al. yields 0.905 mAP and 0.874 F1. The performance gap suggests that global spatial interaction is useful for visibility classification, particularly because occlusion status may depend on evidence distributed across the hand and its surrounding context rather than on local appearance alone.

Calibration and Qualitative Behavior

The F1 score is relatively insensitive to the binarization threshold. Performance peaks near 0.5 and remains above 0.88 between thresholds of approximately 0.3 and 0.7. This stability is favorable for deployment because it reduces the need for precise threshold calibration across operating conditions.

Figure 3

Figure 3: F1 performance remains stable across a broad range of visibility-score thresholds and peaks near 0.5.

The qualitative examples cover self-occlusion, image truncation, and occlusion by external objects. The proposed model produces the most reliable visibility assignments across these categories. In particular, it appears to distinguish a joint that is geometrically inferable from a joint that is directly observed, which is the conceptual distinction motivating the task.

Figure 4

Figure 4: Qualitative predictions for self-occlusion, image truncation, and object-induced occlusion.

Nevertheless, the evaluation remains primarily classification-oriented. The paper does not report calibration metrics such as expected calibration error, Brier score, reliability diagrams, or selective risk curves. Since the predicted values are subsequently used as triangulation weights, calibration quality may matter independently of mAP and F1. A predictor can rank joints correctly while producing scores whose magnitudes are not proportional to measurement reliability.

Visibility-Weighted Multi-View Triangulation

The downstream experiment evaluates whether visibility estimates improve automatic 3D hand annotation. The authors triangulate WiLoR 2D keypoints using DLT across three multi-view datasets: DexYCB, HO3D, and H2O. They compare unweighted triangulation, triangulation weighted by WiLoRโ€™s per-view detector confidence, and triangulation weighted by the proposed per-joint visibility scores.

The distinction between detector confidence and joint visibility is important. A hand detector may assign a high confidence to an entire hand crop even when a particular finger or joint is severely occluded. Per-joint visibility provides a more granular reliability estimate and can suppress corrupted 2D keypoints selectively.

Visibility-weighted triangulation achieves lower median, mean, and interquartile-range reprojection errors on all three datasets. The largest gain occurs on HO3D, which has fewer views and substantial hand-object occlusion. There, the mean reprojection error decreases by 10.1%. This result supports the practical value of modeling visibility independently from pose coordinates: views containing unreliable observations of a particular joint contribute less to its 3D reconstruction.

Figure 5

Figure 5: Visibility weighting reduces triangulation reprojection errors across DexYCB, HO3D, and H2O.

The qualitative triangulation results show that weighting views according to joint visibility suppresses 2D estimates displaced by occlusion and improves the consistency of 3D joint reprojections.

Figure 6

Figure 6: Per-joint visibility weighting improves triangulated 3D locations in regions affected by occlusion and erroneous 2D detections.

The downstream experiment also exposes a limitation. The weighting function is treated as a direct operational use of the predicted visibility probability, but the paper does not investigate alternative robust estimators or weight transformations. For example, reliability-aware M-estimation, visibility-conditioned covariance prediction, or joint-specific uncertainty models could exploit the continuous scores more effectively than direct weighted DLT. The observed improvements nevertheless establish that even an uncomplicated weighting strategy produces measurable benefits.

Theoretical and Practical Implications

Theoretically, the paper separates observability from localization. This separation is valuable because pose estimators often infer occluded joints through learned kinematic priors, whereas downstream systems may need to know whether a prediction is supported by image evidence. A visibility detector therefore supplies a semantic uncertainty variable that is not recoverable from joint coordinates alone.

The results also provide evidence for task-relevant transfer from structured pretrained models. HaMeR and WiLoR outperform generic image representations despite similar or larger parameter counts elsewhere in the comparison. The result suggests that pretraining objectives encoding articulated geometry can transfer to auxiliary perceptual attributes such as visibility, even when those attributes were not explicitly supervised during pretraining.

Practically, the released package can be integrated into multi-view annotation pipelines, HPE confidence estimation, occlusion-aware tracking, and robotic manipulation systems. In AR/VR, per-joint visibility can be used to gate interaction logic or avoid treating hallucinated fingertip positions as direct observations. In dataset construction, it can identify views that should receive reduced influence during triangulation or manual-review prioritization.

The method also has implications for video-based inference. Temporal modeling could enforce visibility consistency, distinguish transient motion blur from persistent occlusion, and exploit visibility transitions as structured events. A video model could jointly predict visibility, keypoint uncertainty, and occlusion boundaries, potentially improving both temporal pose stability and automatic annotation.

Limitations and Future Directions

The study has several limitations. First, HInt provides the primary supervision and evaluation domain, so broader cross-dataset generalization remains insufficiently characterized. The downstream datasets demonstrate utility, but they are used mainly to evaluate triangulation rather than direct visibility classification with independent visibility labels.

Second, the definition of visibility merges physical occlusion and image truncation into a single negative class. These phenomena have different causes and may require different responses in downstream systems. A richer label space could distinguish self-occlusion, object occlusion, inter-hand occlusion, truncation, blur, and low-resolution ambiguity.

Third, the model relies on hand crops. Errors in the upstream detector may affect both the visible evidence and the coordinate frame, yet detector robustness is not analyzed systematically. Joint visibility prediction under imperfect or partial crops would be important for end-to-end deployment.

Finally, the reported fine-tuning degradation motivates but does not settle the question of adaptation. Future work should test parameter-efficient fine-tuning, multi-task learning with pose and visibility, uncertainty calibration, temporal architectures, and explicit probabilistic triangulation. A particularly useful extension would predict a per-joint observation covariance rather than a scalar visibility probability, enabling geometry-aware fusion across cameras.

Conclusion

โ€œHand Visibility Detector: Per-Keypoint Visibility Estimation for Handsโ€ (2608.11574) establishes per-joint hand visibility estimation as a standalone task and demonstrates that frozen, hand-specific HPE backbones provide an effective basis for it. With only a 0.83M-parameter trainable head, the method achieves 0.931 mAP and 0.896 F1 on HInt, outperforming ImageNet-pretrained visibility baselines by 3.4 mAP points. Its downstream use in visibility-weighted triangulation reduces mean reprojection error by up to 10.1% across three multi-view datasets. The findings support explicit observability modeling as a useful complement to pose estimation, while leaving open important questions concerning calibration, cross-domain generalization, temporal consistency, and uncertainty-aware 3D reconstruction.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper introduces a computer vision system called Hand Visibility Detector.

Computer vision means teaching computers to understand pictures and videos. In this case, the computer looks at a picture of a hand and tries to identify its 21 important points, called keypoints. These points include finger joints, fingertips, and the wrist.

The system does not only find where these points are. It also estimates whether each point is:

  • Visible, meaning the camera can directly see it.
  • Invisible, meaning it is covered by something, hidden behind another finger, or outside the picture.

Knowing which joints are visible is useful because a computerโ€™s estimate is usually less trustworthy when a joint is hidden.

2. What questions did the researchers ask?

The researchers focused on several main questions:

  1. Can visibility be predicted for every hand joint separately? For example, can the system tell that one fingertip is visible while another finger joint is hidden?
  2. Can knowledge from existing hand-pose systems improve visibility detection? Large hand-pose models have already studied millions of hand images. The researchers wanted to see whether this learned knowledge could help with the new task of detecting visibility.
  3. Does knowing visibility improve 3D hand reconstruction? If cameras view a hand from several angles, can visibility information help combine those views into a more accurate 3D hand?

The paper is important because earlier systems often used visibility only as a small supporting feature. This paper studies it directly as its own task.

3. How did the researchers do it?

The basic idea

The researchers started with a powerful hand-pose model that had already been trained on a very large number of hand images. This model acts like an experienced observer who already understands hand shapes and finger positions.

They then added a small extra part called a visibility head. Its job is to look at the information produced by the original model and give each of the 21 joints a score between 0 and 1:

  • A score near 1 means the joint is probably visible.
  • A score near 0 means the joint is probably invisible.

For example, a score of 0.9 might mean โ€œvery likely visible,โ€ while 0.1 might mean โ€œprobably hidden.โ€

The original hand-pose model was kept frozen, meaning its knowledge was not changed. Only the small visibility head was trained. This is similar to giving an expert a new checklist instead of retraining the expert from the beginning.

The data

The researchers trained and tested the system using the HInt dataset. This dataset contains more than 30,000 images and video frames from varied, real-world situations, including:

  • Images found online
  • First-person videos
  • Hands partly hidden by objects
  • Hands cut off by the edge of an image
  • Fingers covering other fingers

People manually labeled whether each hand joint was visible. These labels served as the correct answers for training and testing.

Measuring performance

The researchers used two main measurements:

  • mAP, or mean average precision: This measures how well the system ranks visible and invisible joints. Higher values are better.
  • F1 score: This balances two types of mistakes: saying a hidden joint is visible and saying a visible joint is hidden. Higher values are better.

They also tested a second use of the system: 3D triangulation.

Triangulation combines the locations of the same joint seen by multiple cameras. Imagine several people pointing at the same spot from different places. If one personโ€™s view is blocked, that personโ€™s guess may be unreliable. The researchers gave less influence to views where a joint was predicted to be invisible.

4. What did they find?

Better visibility predictions

The proposed system performed better than the two comparison systems.

Method mAP F1 score
Earlier method 1 0.895 0.858
Earlier method 2 0.897 0.860
Hand Visibility Detector 0.931 0.896

These results show that a hand-specific model, already trained to understand hand poses, is especially useful for predicting visibility.

The system worked well in different situations, including:

  • One finger hiding another part of the hand
  • A hand being partly outside the image
  • Objects blocking parts of the hand

The pretrained hand model mattered

The researchers compared different types of computer models. The best results came from HaMeR and WiLoR, which were specifically trained to understand hands.

A general image model called DINOv3 performed fairly well, but it was not as accurate. This suggests that knowing about hand structure is more useful than simply being good at recognizing images in general.

Interestingly, changing the large hand model through fine-tuning made the results much worse. Fine-tuning means continuing to train an already trained model for a new task. In this case, it damaged some of the useful hand knowledge. Keeping the original model frozen worked better.

The attention mechanism helped

The visibility head included a tool called a Gated Attention Unit, or GAU. Attention allows the computer to compare different parts of an image.

This matters because deciding whether a joint is visible may require looking at the whole hand. For example, the computer might notice that a fingertip is hidden because another finger is in front of it.

When the researchers removed this attention component, performance dropped. This showed that understanding relationships between different image locations improved visibility prediction.

More accurate 3D hand reconstruction

The researchers also tested whether visibility scores could improve 3D hand reconstruction from several camera views.

They compared:

  1. Treating every camera view equally
  2. Weighting views using the hand detectorโ€™s confidence
  3. Weighting each joint using its individual visibility score

The third approach worked best on all three tested datasets: DexYCB, HO3D, and H2O.

The largest improvement occurred on the HO3D dataset, where many hands were hidden by objects and there were fewer camera views. The average image error decreased by 10.1%.

In simple terms, the system learned to โ€œlisten moreโ€ to camera views where a particular joint was clearly visible and โ€œlisten lessโ€ to views where that joint was blocked.

5. Why are these results important?

A hand-pose system may make a prediction even when part of the hand is hidden. Without a visibility estimate, users may not know which parts of the prediction can be trusted.

The Hand Visibility Detector adds this missing information. It can help computers:

  • Build more accurate 3D hand models
  • Improve augmented and virtual reality experiences
  • Understand hand movements in videos
  • Help robots interpret how people hold and use objects
  • Create better training data for future hand-recognition systems

The method is also relatively cheap to train because only a small additional part of the model needs to learn. The paper reports that this extra part contains only about 0.13% of the total modelโ€™s parameters.

Conclusion

The paper shows that predicting whether each hand joint is visible should be treated as an important task on its own, not merely as a small addition to hand-pose estimation.

By adding a small visibility detector to an already powerful hand model, the researchers achieved better predictions than earlier methods. They also showed that these predictions can make 3D hand reconstruction more accurate, especially when objects or other fingers hide parts of the hand.

In the future, the system could be improved to work with video over time, making its visibility predictions steadier from one frame to the next.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • The definition of visibility conflates physical occlusion with image truncation by labeling joints outside the frame as invisible; it remains unclear whether the model can distinguish these distinct phenomena.
  • The reliability, consistency, and inter-annotator agreement of the HInt visibility labels are not analyzed, particularly for partially visible, blurred, or ambiguously occluded joints.
  • Evaluation is conducted primarily on HInt, leaving cross-dataset generalization to unrelated domains, camera systems, hand scales, demographics, and environmental conditions unresolved.
  • The reported HInt evaluation uses ground-truth hand bounding boxes, so performance with inaccurate, missed, tightly cropped, or overlapping detector boxes is not established for the standalone visibility task.
  • The model is evaluated on cropped single-hand images, leaving its behavior in full images, two-hand interactions, handโ€“object interactions, and severe hand overlap unexplored.
  • The paper does not report performance separately for self-occlusion, object occlusion, occlusion by another hand, truncation, motion blur, low resolution, or varying occlusion severity.
  • It is unclear whether the detector estimates true image visibility or exploits correlations between HPE backbone features and likely joint locations; experiments with synthetic occlusions, counterfactual occluders, or controlled visibility changes are needed to establish causal sensitivity to occlusion.
  • The model outputs probabilities, but probability calibration is not evaluated; mAP and F1 do not reveal whether a score of, for example, 0.8 corresponds to an approximately 80% likelihood of visibility.
  • The fixed threshold of 0.5 for F1 is not compared with thresholds selected on a validation set or with operating points optimized for downstream annotation, safety, or selective prediction.
  • Performance is averaged across joints, masking potentially important differences between wrist, palm, and fingertip visibility; per-joint metrics and confusion patterns are not reported.
  • The experiments do not assess uncertainty quality, abstention, or riskโ€“coverage behavior, which are important if visibility scores are used to determine whether a joint should be trusted.
  • The claim that frozen pretrained features are preferable is based on a limited fine-tuning comparison; alternative strategies such as partial fine-tuning, adapters, low-rank updates, layer-wise learning rates, or larger visibility-specific training sets are not investigated.
  • The severe degradation of the fine-tuned WiLoR model is not explained; it is unclear whether it results from optimization settings, overfitting, catastrophic forgetting, class imbalance, or a mismatch between the fine-tuning objective and the frozen-backbone setup.
  • The contribution of the pretrained hand prior is not isolated from model scale, input preprocessing, training data, or architecture differences through a fully controlled pretraining and capacity study.
  • The architecture ablation does not examine key design choices such as head width, pooling strategy, attention type, spatial resolution, loss function, class weighting, or joint-to-joint graph modeling.
  • Binary cross-entropy assumes hard binary labels and independent joint targets, leaving partial visibility, annotator uncertainty, and structured dependencies between neighboring joints unmodeled.
  • The training protocol does not discuss class imbalance across visible and invisible joints or provide per-class precision, recall, and average precision, making failure modes difficult to diagnose.
  • The comparison with prior visibility estimators may not be fully controlled for input resolution, crop generation, augmentation, backbone pretraining, parameter count, or implementation details, limiting the strength of the baseline comparison.
  • The study does not compare against strong modern alternatives such as visibility heads attached to other large hand models, multimodal models, heatmap-based visibility predictors, or models trained with explicit occlusion augmentation.
  • The downstream triangulation experiment evaluates reprojection error rather than the accuracy of the reconstructed 3D joints against reliable 3D ground truth; lower reprojection error does not necessarily imply better 3D annotation.
  • The triangulation evaluation depends on WiLoR 2D keypoints and detector boxes, so the independent contribution of visibility estimation is not separated from errors or biases in the chosen HPE system.
  • The paper does not test whether visibility weighting improves downstream 3D pose, mesh, temporal tracking, or task-level performance beyond reprojection error.
  • The visibility-weighting formulation is insufficiently specified and is not compared with alternative weighting functions, learned fusion, robust estimators, confidenceโ€“visibility combinations, or joint-specific weighting schemes.
  • The triangulation results are not broken down by number of views, camera baseline, calibration quality, occlusion level, joint type, or detector failure mode, leaving the conditions under which weighting is most beneficial unclear.
  • The methodโ€™s robustness to camera calibration errors, asynchronous views, missing views, inconsistent detections, and severe multi-view outliers is not evaluated.
  • The paper does not investigate whether visibility estimates from one view are correlated with 2D keypoint confidence or whether combining both signals provides additional benefit over either signal alone.
  • Temporal consistency is identified as future work, but no video model, temporal smoothing baseline, latency analysis, or evaluation of flickering visibility predictions is provided.
  • Runtime, memory use, and latency for practical inference are not reported, despite the backbone containing approximately 631 million parameters; the claim of low-cost training does not establish low-cost deployment.
  • The released package is described as ready to use, but reproducibility across hardware, software versions, pretrained checkpoints, and real-world input conditions is not empirically validated.
  • The dataset and experiments do not establish fairness or robustness across skin tones, hand shapes, ages, left versus right hands, accessories, or culturally diverse interaction contexts.
  • The paper does not examine adversarial or naturally difficult cases in which visual evidence contradicts the learned hand prior, such as unusual poses, malformed hands, strong reflections, gloves, tattoos, or hand-like objects.
  • The practical decision boundary between โ€œvisibleโ€ and โ€œreliably localizableโ€ remains unresolved: a joint may be visually observable but still too blurred, small, or ambiguous for accurate 2D or 3D annotation.

Practical Applications

Immediate Applications

The released model and demonstrated pipeline support the following applications with current technology, provided that input images contain detectable hands and the deployment environment is reasonably similar to the training data.

  • Occlusion-aware AR/VR hand tracking โ€” AR/VR, gaming, and HCI. Integrate the per-joint visibility probabilities with an existing hand-pose estimator to distinguish directly observed joints from joints inferred through occlusion. An AR/VR system could reduce the influence of unreliable finger joints, suppress visual jitter, and avoid placing virtual interactions on hand regions that are hidden by objects or the other hand. Dependencies: The system requires a hand detector or crop generator, such as WiLoR, and sufficient GPU or edge-device capacity. Visibility scores should be calibrated for the target headset camera, lighting, and user population.
  • Reliability-aware gesture interfaces โ€” consumer software and accessibility technology. Use joint visibility as a gating signal for gesture recognition. For example, a sign-language, touchless-control, or presentation interface could accept a gesture only when the key joints relevant to that gesture are visible, or request the user to reposition their hand when confidence is low. Dependencies: Visibility does not indicate whether the estimated joint position is geometrically accurate; it only distinguishes visible from occluded or out-of-frame joints. A gesture recognizer must therefore combine visibility with pose confidence and temporal information.
  • Visibility-weighted multi-camera 3D hand reconstruction โ€” robotics, motion capture, and computer vision. Apply the paperโ€™s visibility-weighted DLT triangulation workflow to multi-view RGB recordings. Each camera view can be weighted separately for each of the 21 MANO joints, reducing the effect of views in which a finger is occluded. This can improve 3D pose reconstruction without requiring marker gloves or specialized depth sensors. Evidence and dependencies: The method reduced reprojection error on DexYCB, HO3D, and H2O, with an improvement of up to 10.1% on HO3D. It assumes calibrated cameras, reliable 2D keypoints, correct hand associations across views, and a valid hand bounding box.
  • Automatic dataset annotation and quality control โ€” academia and industrial AI development. Researchers can use visibility scores to automatically flag problematic frames, reject low-quality keypoints, and produce more reliable 3D hand annotations. In annotation workflows, views with invisible joints can be down-weighted rather than treated equally, reducing manual correction effort for datasets involving grasping, manipulation, or egocentric interaction. Dependencies: Human review remains important for borderline cases and unusual occluders. The reported results are based on HInt and three multi-view datasets, so performance should be validated on new cameras, domains, and annotation conventions.
  • Occlusion-aware robot perception โ€” industrial robotics and collaborative robots. A robot observing a personโ€™s hand while they manipulate tools or objects can use per-joint visibility to identify which portions of the hand are observable. The robot may then avoid using hidden joints for immediate pose control, select a different camera viewpoint, or wait for a clearer observation before executing a dexterous action. Dependencies: Safety-critical robots should not treat the visibility probability as a certified measurement. The system needs independent uncertainty estimation, low-latency inference, hand-object tracking, and validation under the robotโ€™s specific operating conditions.
  • Improved hand-object interaction analysis โ€” manufacturing, logistics, and ergonomics. In video analysis of assembly, picking, packaging, or tool use, visibility-weighted pose estimates can prevent occluded fingers from creating false contact points or erroneous grasp geometries. This can improve activity recognition, ergonomic assessment, and process monitoring. Dependencies: The method does not itself estimate contact, object pose, or hand identity. These capabilities require integration with object detection, tracking, and interaction-recognition models.
  • Interactive annotation tools for hand videos and images โ€” research software and media production. A labeling interface can display a visibility score for every joint, automatically mark joints as likely occluded or out of frame, and prioritize ambiguous frames for human review. This can accelerate annotation for sign-language datasets, animation references, sports analysis, and human-robot interaction recordings. Dependencies: The modelโ€™s definition of โ€œinvisibleโ€ includes both occluded and outside-the-frame joints, so annotation tools should expose or separately infer these two cases when that distinction matters.
  • Camera-view selection in multi-camera capture systems โ€” motion-capture studios and teleoperation. A capture controller can select the camera with the highest visibility for each joint or trigger camera repositioning when critical joints become hidden. This is particularly useful for teleoperation and dexterous manipulation, where thumb and fingertip visibility may be more important than visibility of the wrist or palm. Dependencies: Real-time control requires benchmarking the 631-million-parameter frozen backbone and visibility head on the intended hardware. A smaller distilled model may be necessary for embedded deployment.
  • Policy and standards for trustworthy pose-based systems โ€” public-sector and organizational governance. The method provides a practical basis for requiring pose systems to report per-joint observability rather than outputting unconditional coordinates. Evaluation protocols for AR, robotics, and gesture recognition could include visibility-aware error reporting and failure cases involving self-occlusion, object occlusion, and image truncation. Dependencies: Visibility probabilities should not be interpreted as universal confidence scores. Standards would need separate calibration tests across demographic groups, camera types, environments, and hand-object configurations.

Long-Term Applications

These applications are plausible extensions of the findings but require temporal modeling, larger-scale validation, hardware optimization, or additional task-specific research.

  • Temporally consistent visibility tracking โ€” video understanding and egocentric computing. Extend the model from individual frames to video by incorporating recurrent, transformer-based, or point-tracking modules. The resulting system could maintain stable visibility states through brief occlusions, distinguish temporary occlusion from permanent out-of-frame disappearance, and reduce frame-to-frame flicker in AR/VR and gesture interfaces. Dependencies: The paper identifies video input as future work. Temporal models require labels for visibility transitions and must avoid propagating an incorrect visibility estimate across many frames.
  • Active perception for dexterous robots โ€” robotics and teleoperation. Robots could use predicted visibility to plan camera motion, hand-over-hand observation, or manipulation trajectories that maximize visibility of task-critical joints. A teleoperatorโ€™s system could automatically choose the best viewpoint or provide feedback when the operatorโ€™s fingers are hidden. Dependencies: This requires differentiable or actionable links between visibility, camera movement, task success, and robot control. Latency, occlusions caused by the robot itself, and safety constraints must also be addressed.
  • Visibility-aware 3D hand and object reconstruction at scale โ€” digital twins and synthetic data. Combining visibility estimates with multi-view reconstruction, MANO fitting, and object tracking could generate large collections of realistic 3D hand-object interactions for training robotics, animation, and computer-vision models. Invisible joints could be inferred from visible structure while being explicitly tagged as uncertain in the generated data. Dependencies: Lower reprojection error does not guarantee accurate 3D geometry. Physical plausibility, camera calibration, hand-object contact constraints, and validation against marker- or glove-based capture would be required.
  • Robust sign-language and fine-grained gesture translation โ€” education, accessibility, and communication. A future recognition system could use visibility masks to select alternative cues when particular fingers are occluded, combine multiple cameras, and provide an interpretable explanation for missed signs. Visibility-aware training could improve recognition in natural settings where hands overlap or interact with objects. Dependencies: Sign-language recognition requires extensive language- and signer-diverse datasets. The current binary visibility labels do not encode finger identity ambiguity, motion blur, linguistic context, or the difference between partial and complete visibility.
  • Uncertainty-aware clinical and rehabilitation monitoring โ€” healthcare. Camera-based rehabilitation systems could use joint visibility to identify when hand-exercise measurements are trustworthy, prompt patients to adjust their hand position, and prevent hidden joints from producing misleading range-of-motion estimates. Similar tools could support remote occupational-therapy monitoring. Dependencies: Clinical deployment requires prospective validation, medical-device compliance, demographic and skin-tone robustness testing, privacy safeguards, and uncertainty measures beyond visibility classification. The paper does not demonstrate clinical accuracy.
  • Humanโ€“robot collaboration with risk-sensitive hand prediction โ€” workplace safety and industrial automation. A future system could combine visibility with hand-pose forecasting to estimate where an occluded hand is likely to move and adjust robot speed or trajectory accordingly. For example, a robot could adopt more conservative behavior when critical joints are hidden behind a tool or workpiece. Dependencies: This requires validated motion forecasting, calibrated probabilistic risk models, and fail-safe behavior. Visibility alone cannot establish that a hand is absent or that a predicted pose is safe.
  • Edge-deployable hand visibility modules โ€” mobile devices, wearables, and embedded systems. Model distillation, quantization, pruning, or a lightweight hand-specific backbone could produce a low-power visibility detector for smartphones, smart glasses, cameras, and wearable devices. Such a product could provide on-device occlusion warnings without transmitting video to the cloud. Dependencies: The current approach freezes a very large backbone, so the reported low training cost does not imply low inference cost. Compression must preserve performance under motion blur, low light, small hands, and diverse camera optics.
  • Visibility-aware learning objectives for future HPE systems โ€” academic research and foundation models. Per-joint visibility could become a standard output of hand-pose and hand-mesh models, enabling downstream systems to separate observed evidence from structural inference. Future models could jointly predict visibility, pose uncertainty, occlusion cause, and temporal consistency. Dependencies: The paper shows that freezing a hand-specific pretrained backbone and learning a lightweight head works well, but it does not establish that visibility estimates are fully calibrated or that the approach generalizes to unseen domains. More datasets with consistent labels and evaluation of calibration, causal occlusion type, and cross-dataset transfer are needed.
  • Privacy-preserving and auditable interaction analytics โ€” policy, education, and workplace systems. Systems could store pose data together with visibility metadata, allowing auditors to determine whether an inferred gesture was based on visible evidence or on model completion. This could support more transparent evaluation of classroom engagement tools, workplace monitoring, and interactive public installations. Dependencies: Such applications require strict data-governance policies and user consent. Visibility metadata improves technical transparency but does not by itself resolve surveillance, biometric privacy, or demographic-bias concerns.

Glossary

  • Amodal segmentation: Segmentation that estimates the complete extent of an object, including portions hidden by occlusion. โ€œHDR~\cite{meng2022hdr} recovers the occluded appearance of the target hand and removes the occluding hand via amodal segmentation and appearance recoveryโ€
  • Auxiliary signal: Additional information used to support a primary task rather than serving as the main prediction target. โ€œvisibility estimation has mainly been used as an auxiliary signal for improving pose estimation.โ€
  • Backbone: The main feature-extraction network on which task-specific layers are built. โ€œOur method consists of (i) a frozen ViT backbone of a pretrained HPE modelโ€
  • Bernoulli distribution: A probability distribution for a binary variable with two possible outcomes. โ€œwhich models per-joint visibility as a Bernoulli distributionโ€
  • Binary cross-entropy loss: A loss function that measures the difference between predicted probabilities and binary target labels. โ€œWe train the head with the binary cross-entropy loss against ground-truth visibility labelsโ€
  • Binarization threshold: A cutoff used to convert a continuous probability into a binary prediction. โ€œF1 score versus binarization threshold.โ€
  • Convolution: A neural-network operation that applies learned local filters to structured data such as images or feature maps. โ€œa 1ร—11 \times 1 convolution first compresses the feature map from CC to dd dimensionsโ€
  • DLT (Direct Linear Transformation): A geometric method for estimating 3D points from corresponding observations across camera views. โ€œWe triangulate 2D keypoints estimated by WiLoR~\cite{potamias2025wilor} in each view using DLT~\cite{ABDELAZIZ2015dlt}โ€
  • Egocentric video: Video recorded from a first-person or body-mounted viewpoint. โ€œincluding web images and egocentric video frames.โ€
  • Feature map: A spatial tensor of learned representations produced by a neural network. โ€œthe encoder produces a feature map FโˆˆRhร—wร—CF \in \mathbb{R}^{h \times w \times C}โ€
  • Fine-tuning: Further training of a pretrained model, typically on a task-specific dataset. โ€œfine-tuning the WiLoR backbone for visibility estimation degrades mAP from 0.931 to 0.622.โ€
  • F1 score: The harmonic mean of precision and recall, commonly used to evaluate binary classification. โ€œWe use mAP (average precision averaged over joints) and F1 scoreโ€
  • Gated Attention Unit (GAU): An attention mechanism that combines self-attention with gated linear units to model dependencies. โ€œGAU is an attention mechanism that integrates self-attention with gated linear unitsโ€
  • Ground-truth: The reference labels or measurements treated as correct for training or evaluation. โ€œWe use ground-truth boxes provided by the datasets for training and evaluationโ€
  • Heatmap enhancement: Improvement of spatial probability maps used to locate keypoints. โ€œutilized its output for visibility-guided 2D joint heatmap enhancement.โ€
  • Hierarchical mixture density network: A neural network that represents predictions with probability mixtures organized at multiple levels. โ€œYe et al.~\cite{Ye_2018_ECCV} proposed a hierarchical mixture density network for hand pose estimation from depth imagesโ€
  • Interquartile range: The range between the 25th and 75th percentiles of a distribution. โ€œVisibility-weighted triangulation achieves lower median, mean, and interquartile rangeโ€
  • Keypoint: A semantically meaningful point on an object, such as a hand joint, used for pose estimation. โ€œthe 21 keypoints of the MANO hand modelโ€
  • Logit: An unnormalized score produced before conversion into a probability, often by a sigmoid or softmax function. โ€œspatial average pooling produces per-joint logitsโ€
  • mAP (mean average precision): The mean of average-precision values, here computed across hand joints. โ€œOur method achieves an mAP of 0.931โ€
  • MANO hand model: A parametric 3D model representing hand shape and pose. โ€œRecent HPE methods typically employ deep learning to estimate the parameters of the parametric MANO hand modelโ€
  • Monocular image: An image captured from a single camera or viewpoint. โ€œlarge-scale pretrained models that achieve highly accurate estimation even from monocular imagesโ€
  • Multimodal Gaussian mixture: A probability model containing multiple Gaussian components, allowing several likely modes or solutions. โ€œa multimodal Gaussian mixture to occluded ones.โ€
  • Occlusion: The obstruction of an object or part of it by another object or by the image boundary. โ€œvisibility under diverse occlusion conditions.โ€
  • Off-the-shelf model: A pretrained, ready-to-use model applied without retraining its main components. โ€œan off-the-shelf HPE model for visualization only.โ€
  • Overfitting: Excessive adaptation to training data that reduces performance on unseen data. โ€œwhile avoiding overfitting to the relatively small visibility-labeled data.โ€
  • Parametric model: A model whose outputs or structure are controlled by a finite set of parameters. โ€œthe parameters of the parametric MANO hand modelโ€
  • Pretraining: Training a model on a large dataset before adapting or applying it to another task. โ€œa pretrained HPE model extracts featuresโ€
  • Prior knowledge: Information learned from previous data or tasks that guides a modelโ€™s predictions. โ€œThis preserves the prior knowledge of the large-scale pretrained modelโ€
  • Reprojection error: The discrepancy between an observed image point and the projection of an estimated 3D point back into the image. โ€œvisibility-weighted triangulation reduces reprojection errorsโ€
  • RANSAC: A robust estimation algorithm that fits a model while identifying and rejecting outlier observations. โ€œeven when outliers are removed with RANSACโ€
  • Self-occlusion: Occlusion in which one part of an object hides another part of the same object. โ€œOur method correctly estimates visibility with the highest confidence in all cases of self-occlusionโ€
  • Sigmoid function: An activation function that maps a real-valued input to a value between zero and one. โ€œa sigmoid function converts them into visibility probabilitiesโ€
  • Spatial average pooling: An operation that averages feature values across spatial locations. โ€œspatial average pooling produces per-joint logitsโ€
  • Spatial dependencies: Relationships between information at different positions in a spatial representation. โ€œthe GAU, which models global dependencies among spatial positionsโ€
  • Triangulation: Estimation of a 3D point from corresponding 2D observations in multiple camera views. โ€œvisibility-weighted triangulation of multi-view 2D keypoints reduces reprojection errorsโ€
  • Unimodal Gaussian distribution: A Gaussian probability distribution with one central mode or peak. โ€œassigns a unimodal Gaussian distribution to visible jointsโ€
  • Vision Transformer (ViT): A transformer-based neural architecture that processes images as sequences of patches or tokens. โ€œWe therefore adopt the pretrained ViT backbones of HaMeR and WiLoR as our Hand Encoder.โ€
  • Visibility head: A task-specific neural-network component that predicts whether individual joints are visible. โ€œOur method consists of (i) a frozen ViT backbone of a pretrained HPE model and (ii) a lightweight visibility head.โ€
  • Visibility-weighted triangulation: Triangulation in which observations are weighted according to estimated joint visibility. โ€œvisibility-weighted triangulation of multi-view 2D keypoints reduces reprojection errorsโ€

Open Problems

We found no open problems mentioned in this paper.

Tweets

Sign up for free to view the 2 tweets with 151 likes about this paper.