---
title: 'Vi-TacMan: Vision-Tactile Manipulation'
url: https://www.emergentmind.com/topics/vi-tacman
type: topic
---

# Vi-TacMan: Vision-Tactile Manipulation

Vi-TacMan is a framework for manipulating previously unseen articulated objects such as doors, drawers, and cabinets by combining coarse, global vision with precise, local tactile feedback. It is deliberately designed to work without explicit kinematic models—no joint types, no joint parameters—relying instead on learned geometric priors and contact regulation. In the formulation introduced in "Vi-TacMan: Articulated Object Manipulation via Vision and Touch," vision proposes grasps and coarse interaction directions from RGB-D observations, while a tactile controller refines those proposals during execution through real-time contact regulation. The central claim is that coarse visual cues are enough when coupled with tactile feedback, and the reported experiments span more than 50,000 simulated and diverse real-world objects with statistically significant gains over baselines, all with \(p<0.0001\) [2510.06339].

## 1. Problem Setting and Conceptual Premise

Autonomous household robots must manipulate articulated objects with widely varying appearance, geometry, and hidden mechanisms, including cabinet doors, drawers, and appliances. In Vi-TacMan, articulated object manipulation means: given an RGB-D observation of an unknown articulated object, autonomously find movable and holdable parts, propose a parallel-jaw grasp, infer a coarse interaction direction, and execute the motion using a tactile-contact controller that regulates contact instead of using explicit joint models [2510.06339].

The framework is motivated by the complementary failure modes of vision-only and touch-only approaches. Vision-centric articulated-manipulation pipelines typically infer joint types, axes, and parameters from RGB-D and then plan actions with that kinematic model. The paper emphasizes that articulation is often hidden, that inference is an ill-posed inverse problem on unseen geometries and categories, and that even state-of-the-art methods trained on large articulation datasets can produce imprecise kinematic estimates that fail when executed on a real robot. Touch-only approaches, particularly Tac/TacMan-style contact-regulating controllers, can manipulate articulated objects without precise kinematic models, but they assume a stable grasp and a coarse interaction direction are already given. Vi-TacMan is presented as the first complete system that systematically exploits this vision–touch complementarity for general articulated manipulation [2510.06339].

A common misconception in articulated manipulation is that reliable execution requires explicit symbolic articulation inference, such as revolute or prismatic joint labels and calibrated joint axes. Vi-TacMan rejects that assumption. Its position is not that kinematics are irrelevant, but that manipulation can succeed without explicitly reconstructing a joint model if vision supplies globally useful hypotheses and touch supplies local corrective feedback. This suggests a shift from model reconstruction to action initialization plus contact stabilization.

## 2. System Architecture and Operational Pipeline

The overall pipeline has two main components. The vision module acts as a high-level proposer. From RGB-D, surface normals, and segmentations, it detects movable and holdable parts, proposes a grasp for a parallel gripper, and predicts a distribution over coarse interaction directions. The tactile controller acts as a low-level executor. Given the grasp and coarse direction, it regulates contact using a GelSight-style tactile sensor, iteratively adjusts the end-effector pose to maintain stable contact while moving in the general direction indicated by vision, and succeeds without any explicit joint model [2510.06339].

The operational sequence begins with movable-part and holdable-part detection, followed by fine-grained segmentation masks for each component. The holdable mask is associated with the movable mask by maximal intersection area. Grasp proposal is restricted to the holdable region. The grasp point \(\boldsymbol{g}\) is the centroid of the holdable mask in 3D; rotations are sampled around this point to find a collision-free parallel-jaw grasp with minimal width, and the one closest to the robot home pose is chosen. The visual module also estimates a distribution over interaction directions \(p(\boldsymbol{d}|\mathcal{V})\), summarized by a mean direction used to seed the tactile controller [2510.06339].

The phrase “no explicit kinematic models” has a precise operational meaning in this system. Vi-TacMan never learns or uses joint type labels, joint axes, or joint parameters. Instead, vision predicts point displacements under small perturbations, uses those displacements to infer an approximate rigid transformation of the movable part, and converts that transformation into a distribution over motion directions rather than a symbolic articulation model. The tactile controller then requires only a direction vector and a stable grasp. This enables manipulation even when real-world kinematics deviate from ideal revolute or prismatic assumptions [2510.06339].

## 3. Geometric and Probabilistic Formulation

A central technical component is the use of surface normals as geometric priors. Each visually observed point is represented as
$$
P = (\boldsymbol{p}, \boldsymbol{c}, \boldsymbol{n}, m, h),
$$
where \(\boldsymbol{p} \in \mathbb{R}^3\) is 3D position in the camera frame, \(\boldsymbol{c} \in [0,255]^3\) is RGB color, \(\boldsymbol{n} \in \mathbb{S}^2\) is the surface normal, \(m \in \mathbb{N}\) is the movable-part label, and \(h \in \mathbb{N}\) is the holdable label. Surface normals are estimated from depth by local finite differences: for each pixel, vectors to the right and below neighbors are taken and their cross product gives the normal. The paper argues that normals constrain plausible motion directions; the reported “Without-normal” ablation drops significantly, while a “Normal-only” baseline based on the Fréchet mean of normals within the movable region already performs surprisingly well [2510.06339].

Instead of predicting a single direction vector, Vi-TacMan models motion direction on the 2-sphere with a von Mises–Fisher distribution:
$$
p(\boldsymbol{d}|\mathcal{V}) = \frac{1}{c(\kappa,\boldsymbol{\mu})}\exp\left(\kappa \boldsymbol{\mu}^\top \boldsymbol{d}\right),\quad \boldsymbol{d}\in\mathbb{S}^2,
$$
where \(\boldsymbol{\mu}\in\mathbb{S}^2\) is the mean direction and \(\kappa>0\) is the concentration. Given samples \(\{\boldsymbol{d}_i\}\), the mean direction is estimated by the Fréchet mean under geodesic distance:
$$
\hat{\boldsymbol{\mu}} = \argmin_{\boldsymbol{\mu}\in\mathbb{S}^2} \sum_{i=1}^{n}\left|\arccos\left(\boldsymbol{\mu}^\top \boldsymbol{d}_i\right)\right|^2.
$$
Because the normalizing constant and \(\kappa\) do not affect the maximizer, the MAP direction is simply \(\boldsymbol{d}^*=\hat{\boldsymbol{\mu}}\). This provides a principled aggregation of noisy direction candidates and an explicit representation of directional uncertainty [2510.06339].

The system decomposes the joint estimation problem into grasp selection and rigid transformation inference:
$$
G^*, T^* = \argmax_{G,T} p(G,T|\mathcal{V})
= \argmax_G p(G|\mathcal{V}) \;\; \argmax_T p(T|\mathcal{V}).
$$
For a grasp point \(\boldsymbol{g}\), a rigid transformation \(T=[R|\boldsymbol{t}]\) induces an interaction direction
$$
\boldsymbol{d} = \frac{(R-I)\boldsymbol{p}+\boldsymbol{t}}{\|(R-I)\boldsymbol{p}+\boldsymbol{t}\|_2},
$$
with \(\boldsymbol{p}\) chosen as the grasp point or relevant contact point. Different points on the same rigidly moving part share the same underlying \(T\) but can yield different \(\boldsymbol{d}\), which explains the coupling between grasp location and direction [2510.06339].

To estimate \(T\), the system predicts per-point displacement vectors \(\boldsymbol{q}_i\in\mathbb{R}^3\) and fits \(T\) to point-displacement pairs via the Kabsch algorithm. The training loss for displacement prediction combines relative magnitude error and directional misalignment:
$$
\mathcal{L} =
\frac{1}{n}\sum_{i=1}^{n}\frac{\|\hat{\boldsymbol{q}}_i-\boldsymbol{q}_i\|_1}{\|\boldsymbol{q}_i\|_1}
+
\frac{1}{n}\sum_{i=1}^{n}\left(1-\frac{\hat{\boldsymbol{q}}_i^\top \boldsymbol{q}_i}{\|\hat{\boldsymbol{q}}_i\|_2\|\boldsymbol{q}_i\|_2}\right).
$$
The first term is a relative \(L_1\) magnitude error, and the second term penalizes directional misalignment through cosine similarity. The intended effect is to learn both how much and in what direction each point moves under articulation [2510.06339].

## 4. Vision Module, Training Data, and Generalization Regime

The vision module processes RGB-D together with derived 3D positions, surface normals, and instance-level masks of movable and holdable parts. Detection is performed with a DINOv3 backbone and a transformer-based head. The optimizer is AdamW with learning rate \(6\times 10^{-6}\) and batch size \(2\). On COCO-style mAP@[0.50:0.95], the detector achieves \(0.86\) mAP, with \(\mathrm{AP}(50)=0.97\) and \(\mathrm{AP}(75)=0.94\). Detector outputs prompt SAM2 to obtain fine-grained masks for movable and holdable parts [2510.06339].

Direction estimation is handled by a PointNet++ displacement estimation network that takes point coordinates \(\boldsymbol{p}_i\), surface normals \(\boldsymbol{n}_i\), and movable-mask flags, and outputs displacement vectors \(\hat{\boldsymbol{q}}_i\). This network is trained with AdamW, learning rate \(1\times 10^{-3}\), and batch size \(32\). The input design reflects the paper’s emphasis that geometry, particularly normals, provides strong articulation priors even without explicit joint supervision [2510.06339].

The simulation training corpus is based on 385 articulated objects from PartNet-Mobility. Train categories are microwaves, refrigerators, storage furniture, and trash cans; test categories are dishwashers, doors, ovens, and tables, establishing category-level generalization. Objects are rendered in SAPIEN with ray tracing from up to 72 viewpoints per object. Movable and holdable labels come from GAPartNet. The final splits are approximately 39.5k training samples, 9.9k validation samples, and 5.8k test samples. The broader simulation environment contains 385 articulated objects from 8 categories and more than 50,000 samples of interactions through different viewpoints and object configurations [2510.06339].

Real-world transfer is studied with four physical objects, each captured from five views with a Femto Bolt depth sensor. Depth is refined by a depth foundation model, yielding relative depth, and a linear model fit via RANSAC maps relative depth to sensor depth for correct scale. The paper attributes generalization to simulation diversity, depth refinement, and the fact that the model predicts coarse action hypotheses rather than full articulated state descriptions. This suggests that sim-to-real transfer is supported less by heavy domain randomization than by task decomposition and representation choice [2510.06339].

## 5. Tactile Control and Contact-Regulated Execution

The tactile subsystem uses a GelSight-style tactile sensor mounted on a parallel gripper. The sensed signal is a high-resolution deformation or texture image of the contact patch, which can be processed to infer local geometry and contact state. Vi-TacMan does not treat touch as a separate perception problem followed by explicit replanning; instead, touch is integrated directly into the control loop that executes the visually proposed motion [2510.06339].

The control loop begins from the grasp \(G\) and coarse direction \(\boldsymbol{d}^*\). The end-effector moves along the coarse direction. At each step, the controller observes a tactile image, computes a contact state \(\mathcal{C}_t\), and solves for an incremental pose update \(T_\Delta \in \mathrm{SE}(3)\):
$$
T_\Delta = \argmin_{T_\Delta \in \mathrm{SE}(3)} f(\mathcal{C}_0,\mathcal{C}_{t+1}),
$$
where \(\mathcal{C}_0\) is a reference stable contact, \(\mathcal{C}_{t+1}\) is the contact after applying \(T_\Delta\), and \(f\) is a metric measuring how different the contact is from the reference. The controller then applies \(T_\Delta\) and repeats until task completion or a stopping criterion. In the paper’s description, this is feedback control in pose space rather than trajectory planning in an articulated state space [2510.06339].

This mechanism is especially important when the initial visual direction is inaccurate. The robot starts moving in the proposed direction, contact deviates from the desired stable pattern, and the controller computes six-degree-of-freedom corrections to restore stability. Over time, the object motion is followed by continually re-optimizing \(T_\Delta\), effectively locking onto the correct trajectory even if the original direction was only approximate. The resulting division of labor is explicit: vision supplies global guidance, and touch supplies local precision [2510.06339].

The practical implication is not that tactile sensing replaces perception altogether, but that it replaces the need for highly accurate visual articulation inference at execution time. This suggests a broader control principle for articulated manipulation in unstructured environments: a policy can be robust if it maintains stable contact and tolerates coarse global priors, even when the mechanism is partially hidden or geometrically atypical.

## 6. Empirical Findings, Limitations, and Broader Visuo-Tactile Context

Direction estimation is evaluated on 5,836 test samples from unseen categories—dishwashers, doors, ovens, and tables—using angular error in degrees between predicted and ground-truth directions. All methods have median errors around \(10^\circ\), which the paper presents as evidence of intrinsic difficulty. Vi-TacMan achieves significantly lower errors, with a narrower distribution and fewer large outliers, than FlowBot3D, Normal-only, and Without-normal. A one-sided paired \(t\)-test reports that all improvements are statistically significant with \(p<0.0001\). In real-world experiments, the hardware setup consists of a Kinova Gen3 7-DoF arm, a parallel gripper with GelSight-like tactile sensor pads, and a Femto Bolt RGB-D sensor. The end-to-end pipeline demonstrates opening cabinet doors, drawers, and related objects without per-object modeling or manual tuning, and the tactile controller can correct for initial direction errors caused by occlusions or sensor noise [2510.06339].

The reported limitations are explicit. The system assumes reasonably accurate RGB-D sensing and a high-resolution GelSight-style tactile sensor with nontrivial fabrication and integration. The evaluated articulation scope is mostly standard doors, drawers, and hinged panels from PartNet-Mobility and similar objects; very complex or compliant mechanisms are not evaluated. Performance depends on prior tactile-control methods and on good contact metrics and image-based registration. PointNet++ processing, displacement estimation at high resolution, and distribution fitting may have nontrivial runtime, although latency is not detailed [2510.06339].

Within the broader visuo-tactile literature, Vi-TacMan occupies a systems-level position. The ViTacTip line of work studies a transparent, compliant fingertip sensor with see-through skin and internal biomimetic pins, together with GAN-based modality conversion between pseudo-visual and pseudo-tactile views. That work is relevant because it offers a concrete hardware and modality-management pattern for visuo-tactile fingers, including texture, pose, and force estimation benchmarks [2402.00199; 2501.02303]. DexViTac extends the visuo-tactile scope to portable, human-centric collection of first-person vision, high-density tactile sensing, end-effector poses, and hand kinematics, and introduces a kinematics-grounded tactile representation aligned with a frozen visual encoder for contact-rich dexterous manipulation [2603.17851]. ViTaMIn-B further develops a bimanual, robot-free interface with compliant visuo-tactile DuoTact sensors, Meta Quest controller tracking, latency-calibrated synchronization, and point-cloud tactile representations for cross-sensor generalization [2511.05858].

These adjacent systems do not alter Vi-TacMan’s core contribution, which is specifically articulated-object manipulation without explicit kinematic models. They do, however, indicate plausible extension paths. A plausible implication is that future Vi-TacMan-style systems could combine its coarse-to-fine articulated manipulation strategy with integrated fingertip visuo-tactile hardware, kinematics-grounded tactile representation learning, or bimanual demonstration interfaces. The authors of Vi-TacMan themselves suggest extending to more complex and less idealized articulated mechanisms, including non-rigid or multi-joint motions, using dexterous hands instead of simple parallel grippers, potentially learning explicit kinematic models online from combined visual–tactile data, and further improving robustness to severe occlusions, clutter, and noisy depth [2510.06339].

Source: https://www.emergentmind.com/topics/vi-tacman