GraspFactory: Automated Robotic Grasping
- GraspFactory is a comprehensive framework that integrates an automated physical apparatus, task-specific modules, and a large-scale 6-DoF grasp dataset.
- Its GRM apparatus automates object resetting, orientation control, and multi-object swapping with precise rotational and translational repeatability.
- The dataset provides over 109 million validated grasps while the task-specific pipeline leverages semantic keypoints and probabilistic models to optimize grasp performance.
Searching arXiv for papers on "GraspFactory" and closely related grasping datasets/apparatus. GraspFactory denotes a set of robotic grasp-data production paradigms centered on the automated generation of grasp trials, grasp candidates, and grasp annotations. In the literature considered here, the term is used in three closely related senses: as a description of the Grasp Reset Mechanism (GRM), a fully automated apparatus for conducting large-scale grasping trials; as a tutorial framing for a task-specific grasp-generation module based on semantic keypoint skeletons and probabilistic object models; and, most explicitly, as a large object-centric dataset containing over 109 million validated 6-DoF parallel-jaw grasps. This suggests a unifying emphasis on repeatable grasp production, standardized representations, and scalable data acquisition for learning and evaluation (DuFrene et al., 2024, Robson et al., 2022, Srinivas et al., 24 Sep 2025).
1. Terminological scope and usage
Within the cited literature, “GraspFactory” does not denote a single artifact. It refers, first, to the operation of the GRM as a fully automated “GraspFactory” that can reset objects, control a 1-D orientation variable, collect data, and swap among multiple objects with no human intervention. It also appears as a tutorial label for building a task-specific grasping module exactly as described by Robson & Sridharan. In the most direct nominal sense, it is the title of a large object-centric grasping dataset built for data-intensive 6-DoF grasp synthesis (DuFrene et al., 2024, Robson et al., 2022, Srinivas et al., 24 Sep 2025).
| Context | What “GraspFactory” denotes | Core emphasis |
|---|---|---|
| GRM | Fully automated apparatus for conducting grasping trials | Reset, orientation control, object swapping, data collection |
| Robson & Sridharan method | Step-by-step task-specific grasping module | Semantic keypoints, pSDF, GMM-based task constraints |
| 2025 dataset | Large object-centric grasping dataset | 109 million+ validated 6-DoF grasps |
A common misconception is that GraspFactory names only the 2025 dataset. The broader record here shows that the term also functions as a description of physical infrastructure and as a workflow label for task-conditioned grasp synthesis. A plausible implication is that the term has become associated less with one implementation than with the general idea of industrialized grasp-data generation.
2. GRM as a physical grasp-production apparatus
The Grasp Reset Mechanism is a two-level fixture whose role is to quickly retract and recenter the test object, rotate it about its vertical axis to any commanded angle , and, in the upper level, swap the object for another one in a magazine of pre-loaded shapes. The lower reset combines a centering cone, a string actuator, and a rotating turntable. The cone is a frustum actuated vertically by a NEMA-stepper driving a miniature ballscrew, with a limit switch marking the fully raised position. A thin nylon string, routed through a pulley and tensioner, attaches to a magnetic insert glued into the bottom of the test object; a second NEMA-stepper rewinds the string to pull the object up onto the cone. Two copper plates at the cone top serve as a home-position switch when the magnetic insert contacts them. After lifting, the cone lowers and deposits the object on a 250 mm-diameter rotating plate spun by a DC stepper with an optical encoder disk and an off-center Hall-effect zero magnet (DuFrene et al., 2024).
The upper reset is a 3-DOF Cartesian arm mounted behind the lower reset, moving along , , and rotating about the same -axis, with each axis implemented by a NEMA-stepper and homing limit switches. Its wrist carries an electromagnetic pick head. During a swap, the lower reset lifts the cone and object, the upper arm swings into position and lowers until its electromagnet couples to the object, the electromagnet is energized so the object drops cleanly from the base insert, the arm stows the old object in a rear magazine and selects a new one, and the arm then returns the new object to the base insert before de-energizing. Together, the lower and upper reset provide four continuous axes of motion plus two discrete coupling states: string engaged or not, and electromagnet on or off.
The GRM’s kinematic description is given by
$T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$
with
The linear -actuator obeys
where is the cumulative motor steps, 0 the steps per screw revolution, and 1 the ballscrew pitch. These relations formalize the apparatus as a deterministic reset-and-placement mechanism rather than an ad hoc experimental jig.
3. Control architecture, sensing, and grasp-trial dataset
The GRM is orchestrated by a formally specified finite state machine,
2
with
3
and events such as 4, 5, 6, 7, 8, 9, and 0. Each event corresponds to a real sensor callback, including a limit switch tripping, encoder count completion, a serial reply from the arm, or rosbag completion. The transition sequence moves from object release and recentering, through turntable rotation, manipulator preparation, grasp execution, post-grasp lift and move-away, optional swap, and termination (DuFrene et al., 2024).
Because the only remaining degree of freedom on the lower reset is rotation about the vertical axis, the system achieves repeatable single-axis orientation control. The sequence is: raise the cone until magnetic coupling center-aligns the object, lower the cone so the object sits on the stainless-steel turntable, rotate the turntable to the commanded angle, then use the Hall-effect zero mark to establish absolute 1 before counting steps to 2. Rotational repeatability was measured as 3 standard deviation over 20 trials, and translational 4 repeatability was 5 mm 6 mm.
During each grasp trial, the system records a timestamped rosbag containing manipulator joint states at 100 Hz, gripper state at 50 Hz, top RGB camera at 30 Hz, side RGB camera at 30 Hz, wrist-mounted RGB-D at 30 Hz, and object pose at 30 Hz via offline detection. A grasp is labeled successful if, after closing the gripper, the reported gap never returns to the fully closed limit: 7 Across 8 trials, the GRM observed 9, giving
0
The corresponding dataset was produced with a Kinova Gen3 robot arm and Robotiq 2F-85 Adaptive Gripper. It includes four canonical shapes—rectangular prism, triangular prism, cylinder, and cone—each equipped with a magnetic socket and top Aruco-marker, except the cone’s colored dot. Trials were conducted over boundary “edge-region” perturbations of the end effector, including 1 translation, 2 translation, and 3, 4, 5 rotations over specified ranges. The prisms were additionally rotated on the turntable to 6 or 7. Each of the 1,020 trials took approximately one minute including planning, motion, data recording, and reset, for a total of approximately 17 hours. The dataset provides manipulator states, object pose, video, and grasp success data for every trial.
4. Task-specific grasp generation as a “GraspFactory” workflow
Robson & Sridharan describe a method for generating robot grasps by jointly considering stability and task- and object-specific constraints. Their step-by-step “GraspFactory” tutorial is built around a three-level representation. Level 1 is class membership, with each object instance assigned to a semantic class from a small number of exemplars. Level 2 is a semantic keypoint skeleton, typically comprising 2–5 manually annotated keypoints on an exemplar; at run time, a stacked-hourglass CNN detects 2D keypoints in each RGB view and multi-view triangulation produces 3D keypoints linked into a skeleton. For each keypoint, a class-specific Euclidean frame is defined using adjacent keypoints. Level 3 is a probabilistic surface representation obtained from a probabilistic signed distance function (pSDF), where each voxel stores a Gaussian 8 and new depth measurements are merged recursively before extracting a uniformly sampled point cloud via marching cubes (Robson et al., 2022).
The relationship between semantic structure and grasp contacts is encoded as a spherical Gaussian Mixture Model anchored at each keypoint frame. For a surface contact point 9 in a keypoint’s Euclidean frame, the method computes a unit vector, scales it to sphere radius 0, converts it to spherical coordinates, and fits a Dirichlet-process GMM over 1 pairs. This yields a likelihood 2 that a new surface direction from keypoint 3 is a good grasp point for the task. At run time, each surface point is scored by the product of the GMM probabilities associated with the nearest link’s two keypoints: 4
The run-time sampling procedure acquires multi-view RGB-D, segments with Detectron2, builds the pSDF, extracts a point cloud with per-point variances, detects and triangulates 3D keypoints, builds the skeleton, computes task scores over surface points, and samples 5 initial candidate contacts by drawing 45 points from the point cloud with probability proportional to 6. For each candidate contact and each gripper yaw angle in 7, the system places a virtual fingertip, aligns the gripper, simulates closing via Eigengrasp, rejects collisions and non-contacts, computes three scoring terms, and selects the top candidate for local optimization in rotation and translation.
The grasp-success model is
8
where 9 is a contact-angle term, $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$0 is a surface-confidence term derived from pSDF uncertainty and local curvature, and $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$1 is a task/object term: $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$2 By tuning $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$3, the method trades off pure stability against task-specific placement. In physical experiments with a Franka Emika Panda and Intel RealSense D415 cameras, six classes and two tasks per class were evaluated over 530 trials. The reported trend was that task-specific models boosted correct task compliance to 90–100% versus 0–83% for the stability-only baseline, while stability remained high at 80–96%, with only small drops when the task forced grasps on narrow or curved handles. For the cup class, the baseline achieved 80% stable grasps, with 83% of those suitable for handover and 12% for pour; the handover model yielded 80% stability with 96% correct handover; the pour model yielded 77% stability with 83% correct pour.
5. GraspFactory as a large object-centric 6-DoF dataset
The 2025 GraspFactory dataset is, to the authors’ knowledge, the largest publicly available object-centric 6-DoF grasping dataset. It contains over 109 million “good” 6-DoF, parallel-jaw grasps collectively for the Franka Panda and Robotiq 2F-85 grippers. The Robotiq subset includes 33,710 distinct watertight triangle meshes sourced from the ABC dataset and 391.38 million collision-free grasp candidates, of which 97.1 million passed physics-based robustness tests. The Franka Panda subset includes 14,690 objects and 227.22 million grasp candidates, of which 12.2 million were deemed feasible under simulation (Srinivas et al., 24 Sep 2025).
All objects are drawn from the ABC dataset, which contains more than 1 million CAD models of mechanical parts. The dataset stores object geometry as watertight triangular meshes in OBJ or PLY format. Each grasp is represented by a successful pose $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$4 and an associated gripper width $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$5 in millimeters: $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$6 For point-cloud-based learning, uniformly sampled 3D point sets, such as 2,048 points per object, are also provided.
To reduce redundancy, the dataset defines a distance metric in $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$7,
$T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$8
where $T_{W\!O}=T_{W\!R}\cdot \underbrace{\begin{bmatrix}R_z(\theta)&t_{cone}\0&1\end{bmatrix}}_{T_{R\!O}(\theta)},$9 are the unit quaternions corresponding to 0. Candidate grasps are generated by sampling a surface point and unit normal, casting three rays within a 1 cone about that normal, finding a second point with nearly antipodal normal, aligning the gripper-finger normals with the line connecting the contact pair, and sampling four equally spaced yaw angles about that axis, each separated by 2. The process is repeated on a mesh decimated by a factor of 0.6 to expose coarser surface patches.
Collision checking discards candidates whose fingers intersect the mesh. The remaining candidates are clustered in 3 and reduced to up to 2,000 poses per object for Panda and 5,000 for Robotiq. Physics evaluation then runs in NVIDIA Isaac Sim using a parallel farm of 2,000 Franka Panda robots or 5,000 Robotiq 2F-85 grippers. Each object is placed in pose 4 with its local 5-axis aligned to world 6, a grasp is executed, and scripted perturbations test force closure and frictional stability. Grasps that maintain contact without slip throughout the perturbation sequence are marked successful. Exported outputs include successful grasps 7 and indices of failed grasps for negative mining.
6. Training, evaluation, limitations, and extensions
For learning, the GraspFactory paper reports baseline training with SE(3)-DiffusionFields. The input is a 2,048-point cloud normalized to unit diameter, the encoder is DGCNN-style, the diffusion network is SE(3)-equivariant, and grasp samples are drawn by annealed Langevin dynamics in pose space. Training on the Franka Panda subset used two NVIDIA RTX-4090 GPUs, batch size 4, and 2,900 epochs, approximately 19 days, with 12,903 objects for training, 1,434 for validation, and 353 held out for test. On held-out GraspFactory objects versus held-out ACRONYM objects, the model trained on GraspFactory achieved 0.75 on self-test and 0.59 on cross-test, whereas the ACRONYM-trained model scored 0.08 on GraspFactory objects. In simulated evaluation on industrial test objects, examples included Hanger with GF 0.97 versus Acronym 0.30 under one seed and 0.93 versus 0.22 under another, and Base with GF 0.75 versus 0.00 and GF 0.72 versus 0.47 (Srinivas et al., 24 Sep 2025).
Real-world hardware evaluation used a UR10e with Robotiq 2F-85 gripper, Segment Anything Model plus CNOS for initial ROI, FoundationPose for CAD-assisted 6D pose estimation 8, and the transform chain
9
Over 30 non-colliding grasps for each of 3 random object poses across 8 parts, for 720 total grasps, average success rates per part ranged from 80% up to 100%. In a separate robustness experiment of 5 grasps over 10 trials for 4 objects, average success was 96%–100%. An additional set of 10 challenging real objects achieved averages from 3/5 to 5/5 successful grasps.
The limitations reported across these works are distinct and complementary. The GRM is limited to single-axis 0 orientation, object size 1 mm and weight 2 kg, or 3 mm and 500 g for magazine swaps, rigid or semi-rigid objects only, and no closed-loop vision correction during grasps. The large dataset assumes uniform mass and friction across all objects, does not incorporate explicit finger shape and contact patch modeling in the learning stage, and was evaluated in the real world with a slightly different gripper and an occlusion-prone perception pipeline. The task-specific grasp-generation work, by contrast, shows that stability-only optimization can be insufficient when a task requires constrained contact placement. This suggests that physical reset mechanisms, semantic task models, and massive object-centric datasets address different bottlenecks in grasping research rather than serving as interchangeable solutions.
Reported extensions likewise follow different trajectories. For the GRM, proposed directions include an 4–5 translation stage under the turntable, wrist force/torque sensing, deformable-object handling via vacuum gripper, multiple GRMs for dataset creation beyond 6 grasps, and real-time vision feedback for active re-centering or collision avoidance. For the dataset, planned extensions include adding the remaining ABC meshes toward approximately 1 million total and 7 grasp candidates, adding other end-effectors such as suction cups and multi-fingered hands, and enriching the data with RGB(D) renderings and textured mesh variants. Taken together, these extensions indicate an ongoing shift from isolated grasp benchmarks toward integrated grasp-production ecosystems.