SymOpt: Dual-Hand Grasp Synthesis
- SymOpt is a pipeline that converts single-hand grasp data into dual-hand grasp datasets by leveraging object and hand symmetries.
- It employs a two-stage process: mirroring initial grasp proposals and refining them via energy-based physical optimization.
- The resulting datasets, enriched with millions of plausible grasps and semantic annotations, significantly advance bimanual manipulation research.
Searching arXiv for "SymOpt" and the cited papers to ground the article. SymOpt is a pipeline for constructing large-scale dual-hand grasp datasets from single-hand grasp data by exploiting object and hand symmetries and then refining the resulting bimanual grasp proposals through physical optimization. It is introduced in the context of "DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions" (Li et al., 26 Sep 2025), where it serves as the data-generation foundation for learning text-guided, affordance-aware dual-hand grasp synthesis. In that formulation, SymOpt addresses a central bottleneck in bimanual manipulation research: existing grasp datasets predominantly focus on single-hand interactions and contain only limited semantic part annotations, whereas robust dual-hand grasp learning requires large-scale, physically plausible, and semantically structured supervision (Li et al., 26 Sep 2025).
1. Problem setting and motivation
Dual-hand object grasping is presented as essential for advanced manipulation tasks, yet large-scale data for such interaction is scarce. Existing datasets are described as focusing mainly on single-hand interactions, with insufficient scale and semantic richness for learning coordinated dual-hand grasps (Li et al., 26 Sep 2025). SymOpt is introduced specifically to leverage the abundance of single-hand grasp datasets and transform them into dual-hand grasp datasets that are physically plausible and scalable.
The pipeline operates on the premise that most everyday objects and the human hands or tooling represented in grasp datasets are approximately bilaterally symmetric. This premise motivates a two-stage construction process: first, dual-hand candidates are synthesized from single-hand grasps through symmetry-based mirroring; second, these candidates are refined through optimization because mirrored configurations do not automatically satisfy physical validity constraints such as non-interpenetration (Li et al., 26 Sep 2025).
This suggests that SymOpt is not merely a data augmentation heuristic. A plausible implication is that it functions as a structured synthesis-and-validation procedure, in which geometric prior structure supplies candidate grasps and optimization enforces physical admissibility.
2. Symmetry-based synthesis from single-hand data
SymOpt begins with a 3D object point cloud and associated single right-hand grasps, with DexGraspNet given as the example source dataset (Li et al., 26 Sep 2025). The first technical step is symmetry plane discovery. The point cloud is centered, and three mirrored versions are created by reflection over the , , and axes. The best symmetry plane is then selected by comparing each mirrored point cloud to the original using the Chamfer Distance:
This procedure operationalizes symmetry detection as a model-selection problem over axis-aligned reflections rather than as unrestricted symmetry estimation (Li et al., 26 Sep 2025). Once the symmetry plane is chosen, the right-hand mesh and its pose parameters , consisting of translation, orientation, and joint angles, are mirrored across the selected plane to generate a left-hand proposal (Li et al., 26 Sep 2025).
At this stage, the mirrored left-hand pose remains parameterized as a right hand, and a proper left-hand configuration is deferred to later processing (Li et al., 26 Sep 2025). The important point is that the mirrored proposal is treated as an initial guess rather than a completed dual-hand grasp. That distinction is central to the architecture of SymOpt: symmetry creates candidate correspondence, but optimization is required to produce physically plausible bimanual grasps.
3. Physical optimization and energy formulation
SymOpt explicitly recognizes that mirrored grasps may interpenetrate the object or the other hand. To correct this, it applies an energy-based physical optimization over the original right-hand pose and the mirrored left-hand pose (Li et al., 26 Sep 2025).
The optimization uses two energy terms. The first penalizes hand-hand interpenetration:
where is 0 if the vertex 1 from the left hand lies inside the right hand, and 2 is its minimal distance to the right-hand mesh (Li et al., 26 Sep 2025).
The second penalizes hand-object interpenetration:
3
where 4 denotes the object mesh vertices (Li et al., 26 Sep 2025).
These terms are combined into the total optimization energy
5
with 6 and 7 as hyperparameters (Li et al., 26 Sep 2025). Adam is used to update 8 and 9 so as to minimize 0, and termination occurs once the energy terms drop below a threshold, indicating that the hands are not penetrating each other or the object (Li et al., 26 Sep 2025).
The pipeline also filters proposals that remain irreconcilable after optimization, such as cases of severe interpenetration (Li et al., 26 Sep 2025). This filtering stage is significant because it means the final dataset is curated not only by synthesis but also by explicit rejection of unsuccessful physical refinements.
4. Dataset construction and scale
The principal output of SymOpt is a large-scale dual-hand grasp dataset. Applied to DexGraspNet, described as an original single-hand dataset with 1.3M grasps over 5,000+ objects, SymOpt produces a dual-hand set with 7.8 million physically plausible grasps over these objects (Li et al., 26 Sep 2025). After selection for diversity and quality, the curated dataset DualHands-Full contains 1.3M grasps for 800+ objects (Li et al., 26 Sep 2025).
A second dataset, DualHands-Sem, is created as a semantically annotated subset. It contains 157 objects from DualHands-Full that are manually segmented into semantic parts, and grasps are labeled with corresponding affordance-based text labels. Only grasps with at least 1 contact on a single part are labeled, with the aim of ensuring clear semantic learning signals (Li et al., 26 Sep 2025).
The following table summarizes the dataset construction reported for SymOpt.
| Dataset | Description | Reported scale |
|---|---|---|
| DualHands-Full | Curated dual-hand grasp dataset selected for diversity and quality | 1.3M grasps for 800+ objects |
| DualHands-Sem | Manually segmented semantic subset with affordance-based text labels | 157 objects |
| Intermediate dual-hand set | Physically plausible dual-hand grasps produced from DexGraspNet by SymOpt | 7.8 million grasps |
The paper further states that SymOpt datasets have an order of magnitude more dual-hand grasps per object than BimanGrasp or other prior approaches (Li et al., 26 Sep 2025). Because the exact comparison values are not fully enumerated in the supplied material, that claim is best understood as a qualitative scale contrast anchored to the reported dataset sizes.
5. Role in DHAGrasp and affordance-aware learning
Within the DHAGrasp framework, SymOpt is the mechanism that supplies the training and evaluation data needed for affordance-aware dual-hand grasp synthesis (Li et al., 26 Sep 2025). The relationship is explicit: the large-scale, physically plausible dual-hand grasp data generated by SymOpt forms the foundation for training the text-guided grasp synthesis model.
DHAGrasp builds on this data using a dual-hand contact representation. For each hand, the representation includes a contact map 2, a part map 3, and affordance directions 4 (Li et al., 26 Sep 2025). The architecture then follows a two-stage design. In Text2Dir, category-level affordance directions 5 are predicted from object shape and text instruction using diffusion modeling; in Dir2Grasp, the object and predicted directions are used to generate full dual-hand grasp parameters (Li et al., 26 Sep 2025).
The training objective for Text2Dir is given as
6
where 7 is the object/text embedding and 8 are noisy directions and masks at time 9 (Li et al., 26 Sep 2025). For Dir2Grasp, the hand translation is normalized by object diameter, 0, and the hand-parameter regression includes
1
(Li et al., 26 Sep 2025). A penetration loss similar to SymOpt’s optimization energy is used to ensure physically plausible generation, and test-time adaptation can further refine generated grasps using energy-based optimization similar to that of SymOpt (Li et al., 26 Sep 2025).
This architecture places SymOpt in a specific methodological position: it is both a dataset-generation pipeline and a source of inductive structure for later synthesis. The same physical plausibility constraints used during dataset creation reappear during generation and test-time adaptation, which suggests a deliberate continuity between data curation and inference-time refinement.
6. Reported performance and comparative significance
The supplied material attributes strong physical quality to data generated by SymOpt. In Isaac Gym physics simulation, the success rate is reported as exceeding 2, in contrast to approximately 3 for previous works such as BimanGrasp (Li et al., 26 Sep 2025). A detailed friction-wise comparison is given:
4
(Li et al., 26 Sep 2025). The same source reports that DualHands-Full contains approximately 1600 grasps per object and that DualHands-Sem contains 40k grasps over 157 objects, with more than 250 grasps per object (Li et al., 26 Sep 2025). DHAGrasp trained on these datasets is described as producing dual-hand grasps that are semantically consistent with text inputs and generalizing to unseen and unsegmented objects (Li et al., 26 Sep 2025).
These results support two distinct claims about SymOpt. First, it increases dataset scale substantially by transforming single-hand data into dual-hand data. Second, the resulting grasps are not merely numerous but physically credible in simulation. A plausible implication is that the optimization stage is essential to both outcomes: without it, mirroring alone would increase nominal quantity but not necessarily usable quality.
7. Conceptual position and related interpretations
SymOpt should be distinguished from several similarly named but methodologically different optimization frameworks. "SimPO: Simultaneous Prediction and Optimization" (Zhang et al., 2022) concerns end-to-end decision-focused learning with a joint weighted loss over predictive and optimization objectives, rather than bimanual grasp dataset construction. "LM-SymOpt" in robotic manipulation (Tang et al., 25 Jan 2025) refers to a language-guided symbolic task planning framework with optimization, where LLMs translate natural language into symbolic plans and trajectory optimization selects among feasible plans. "SOSOPT" (Seiler, 2013) is a Matlab toolbox for sum-of-squares polynomial optimization. These works share optimization-oriented nomenclature but address different technical problems.
By contrast, SymOpt in (Li et al., 26 Sep 2025) is specifically a pipeline for converting single-hand grasp corpora into large-scale dual-hand grasp datasets using symmetry and optimization. Its defining components are axis-reflection-based symmetry plane selection, mirrored hand-pose construction, interpenetration-aware energy minimization, and dataset curation through optimization success and diversity filtering (Li et al., 26 Sep 2025).
A common misunderstanding would be to interpret SymOpt as a purely symbolic or purely semantic framework because of its name. The supplied evidence indicates the opposite: its core mechanism is geometric and physical. Semantics enters later through DualHands-Sem and the DHAGrasp training pipeline, where affordance-based labels and text guidance are layered on top of physically plausible dual-hand grasp data (Li et al., 26 Sep 2025). Another possible misunderstanding is to treat symmetry as sufficient for bimanual grasp generation. The optimization stage directly contradicts that simplification, because mirrored grasps are explicitly said not to guarantee physical plausibility (Li et al., 26 Sep 2025).
In the context provided, SymOpt’s main contribution is therefore infrastructural. It bridges the gap between abundant single-hand data and the supervision demands of affordance-aware dual-hand grasp learning, while preserving physical plausibility through explicit optimization. That role makes it a dataset-construction methodology with downstream consequences for generalization, semantic alignment, and robust bimanual manipulation (Li et al., 26 Sep 2025).