Papers
Topics
Authors
Recent
Search
2000 character limit reached

GRASP-1k: Robotic Grasping & Geospatial Benchmark

Updated 9 July 2026
  • GRASP-1k is a dual-usage dataset that names both a robotic grasping corpus and a geospatial reasoning benchmark with structured outputs.
  • In robotics, it provides 1,020 trials with detailed sensor data, controlled perturbations, and quantifiable success measures for grasp planning.
  • For geospatial reasoning, it offers 1,071 segmentation examples paired with language instructions to evaluate reinforcement-learning policies on OOD imagery.

GRASP-1k denotes two distinct research datasets introduced in different arXiv works: a robotic grasping dataset associated with the Grasp Reset Mechanism (GRM) in "The Grasp Reset Mechanism: An Automated Apparatus for Conducting Grasping Trials" (DuFrene et al., 2024), and a geospatial reasoning-and-segmentation benchmark introduced in "GRASP: Geospatial pixel Reasoning viA Structured Policy learning" (Jiang et al., 23 Aug 2025). In the first usage, GRASP-1k is a corpus of 1,020 physical grasp trials collected with a fully automated robotic apparatus; in the second, it is an out-of-domain benchmark of 1,071 reasoning-segmentation examples for natural-language-guided geospatial mask generation. The shared name reflects different senses of “grasp”: physical manipulation in robotics and structured grounding in spatially localized vision-language reasoning.

1. Nomenclature and scope

The term GRASP-1k is not unique to a single benchmark. In the available literature, it names two separate datasets with different modalities, objectives, and evaluation regimes.

Usage Paper Core content
Robotic grasping (DuFrene et al., 2024) 1,020 physical grasp trials collected with the GRM
Geospatial pixel reasoning (Jiang et al., 23 Aug 2025) 1,071 reasoning-segmentation examples for OOD evaluation

In the robotic setting, the dataset is explicitly tied to a hardware-software platform for automated resetting, object swapping, and repeatable grasp execution (DuFrene et al., 2024). In the geospatial setting, the benchmark supports reinforcement learning for segmentation from language instructions, with structured outputs consisting of bounding boxes and points rather than direct mask supervision (Jiang et al., 23 Aug 2025).

This suggests that GRASP-1k is best understood as a context-dependent dataset name whose meaning must be resolved from the associated paper and task domain.

2. GRASP-1k in robotic grasping research

In (DuFrene et al., 2024), GRASP-1k is a dataset of 1,020 grasps created with a Kinova Gen3 robot arm and Robotiq 2F-85 Adaptive Gripper. The stated purpose is to enable training of learning models and to demonstrate the capabilities of the GRM. The dataset includes ranges of grasps conducted across four objects and a variety of orientations, with manipulator states, object pose, video, and grasp success data provided for every trial.

The overall statistics reported are 1,020 total trials and 715 successful grasps, approximately 70% (DuFrene et al., 2024). The objects and orientations are specified as follows: a rectangular prism of 40×40×10540\times 40\times 105 mm with θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}; a triangular prism of 50×50×10550\times 50\times 105 mm with θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}; a cylinder of 40×105\varnothing 40\times 105 mm with θ={0 only}\theta=\{0^\circ\text{ only}\}; and a cone with base 40\varnothing 40, height $105$ mm, again with θ={0 only}\theta=\{0^\circ\text{ only}\}. All shapes are rigid, with a magnetic base insert and an on-top ArUco marker, except that the cone uses a colored dot (DuFrene et al., 2024).

For each (object,θ)(\text{object},\theta) combination, 15 equally spaced deviations were applied in one degree of freedom of the gripper. The perturbations include translations θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}0 and θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}1, as well as rotations in roll about θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}2 over θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}3, pitch about θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}4 over θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}5, and yaw about θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}6 over θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}7. Two grasp types are represented: top-down and side-approach (DuFrene et al., 2024).

The recorded modalities include manipulator joint states, comprising positions and velocities at 100 Hz; gripper finger widths and forces; object pose from an ArUco tracker or color-dot tracker at 30 Hz; top-view RGB at 30 Hz; side-view RGB at 30 Hz; a wrist-mounted Intel RealSense RGB-D stream at 30 Hz; and a binary success label per trial (DuFrene et al., 2024).

The paper also reports sample statistical highlights. For top grasps under X-translation, success rates are 68% for the rectangle, 35% for the triangle, 53% for the cylinder, and 40% for the cone. For top grasps under Y-rotation, the rates are 95%, 90%, 47%, and 33%, respectively. For side grasps under X-rotation, the corresponding rates are 100%, 97%, 93%, and 100% (DuFrene et al., 2024).

3. GRM apparatus, reset pipeline, and software interface

The robotic GRASP-1k dataset is inseparable from the Grasp Reset Mechanism. The GRM is described as a fully automated apparatus for conducting large-scale grasping trials, automating the process of resetting a grasping environment, repeatably placing an object in a fixed location and controllable 1-D orientation, collecting data, and swapping between multiple objects with no human intervention (DuFrene et al., 2024).

Its lower reset subsystem consists of a centering cone driven by a vertical ballscrew and NEMA-17 stepper, a retractable string driven by a second NEMA-17 stepper with a gravity-loaded tensioner, and a rotating platform of 25 cm diameter driven by a DC motor with a quadrature-style encoder and a Hall-effect “zero” sensor (DuFrene et al., 2024). Electrical home-position sensing is implemented with a ballscrew limit switch for cone-up detection and copper-plate contacts around the cone that detect when the magnetic insert on the object draws current, indicating that the object is seated and centered (DuFrene et al., 2024).

The upper reset, or object swapping subsystem, is a back-mounted 3-DOF mini-gantry with two linear NEMA-17 steppers for θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}8 and θ{0,15,30,45}\theta\in\{0^\circ,15^\circ,30^\circ,45^\circ\}9, one rotary stepper about 50×50×10550\times 50\times 1050, and an electromagnet at its wrist. The described procedure is: the lower reset lifts the object on its cone, the upper arm swings in and powers the electromagnet, lifts the object off the string insert, places it in a storage bin, picks up the next pre-loaded object by magnet, and returns it to the cone (DuFrene et al., 2024).

The control architecture is a distributed ROS-based system tested under Melodic and Noetic, with two logical hosts. A master control computer runs a FlexBE hierarchical state machine, arm action servers and clients, and data-collection nodes, while an on-GRM Raspberry Pi with two Arduinos handles stepper drivers, an H-bridge, limit-switch interrupts, and electromagnet control (DuFrene et al., 2024).

The high-level behavior is decomposed into FlexBE states: TestControl, TrialControl, ResetEnvironment, ExecuteGrasp, DataCollection, EvaluateSuccess, and LogResults (DuFrene et al., 2024). Example ROS interfaces include /grm/lower_reset with goal {angle: θ}, /grm/swap_object with goal {next_object_id}, /arm/execute_grasp with goal {SE(3) pose}, and /data_recorder/start_stop for rosbag control (DuFrene et al., 2024). The open-source materials include CAD and a bill of materials, and the software interface exposes ROS action servers such as grm_lower_reset.action, grm_rotate.action, grm_swap_object.action, arm_execute_grasp.action, and arm_set_home.action (DuFrene et al., 2024).

The reset precision is quantified in two ways. The centering cone gives an accurate 2-D XY home location with 50×50×10550\times 50\times 1051 mm 50×50×10550\times 50\times 1052, and the platform rotates to target 50×50×10550\times 50\times 1053 about its vertical axis with 50×50×10550\times 50\times 1054 (DuFrene et al., 2024). The paper also reports pose repeatability from 20 resets tracked with ArUco as 50×50×10550\times 50\times 1055 mm and 50×50×10550\times 50\times 1056 (DuFrene et al., 2024).

4. Formal task definition and evaluation in the robotic dataset

The robotic paper specifies a grasp representation

50×50×10550\times 50\times 1057

where

50×50×10550\times 50\times 1058

and

50×50×10550\times 50\times 1059

Within the dataset, only θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}0 around world-θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}1 is varied via the GRM, while θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}2 and θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}3 remain at zero. End-effector perturbations add θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}4 along one axis or one Euler angle (DuFrene et al., 2024).

The success criterion is defined during a lift-and-move to a fixed target 25 cm away. A grasp is labeled successful, θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}5, if and only if two conditions both hold: the gripper never fully closes, meaning the width never equals zero, and the final object centroid at the target lies within θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}6 mm. Formally,

θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}7

The evaluation metrics include success rate per θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}8, transition boundary θ{0,20,40,60}\theta\in\{0^\circ,20^\circ,40^\circ,60^\circ\}9 or 40×105\varnothing 40\times 1050 where success transitions to failure, and pose repeatability (DuFrene et al., 2024).

The usage notes identify several experimental roles for GRASP-1k. These include supervised learning, such as training a CNN to predict 40×105\varnothing 40\times 1051 from RGB/D and pose; boundary estimation, namely learning 40×105\varnothing 40\times 1052 as a function of object shape and 40×105\varnothing 40\times 1053; domain adaptation, specifically sim2real bootstrapping using fine angle increments; and benchmarking new grasp planners on the same 1,020-trial protocol (DuFrene et al., 2024). A plausible implication is that the dataset was designed not only for aggregate success-rate reporting, but also for sensitivity analysis around failure boundaries and controlled physical repeatability.

5. GRASP-1k in geospatial pixel reasoning

In (Jiang et al., 23 Aug 2025), GRASP-1k is a benchmark for geospatial pixel reasoning, a task defined as generating segmentation masks directly from natural-language instructions. The paper describes prevailing multimodal large-language-model systems as co-training a LLM and a mask decoder with dense pixel supervision, then introduces GRASP as a structured policy-learning framework in which a multimodal LLM emits task-relevant bounding boxes and positive points from a vision-language instruction, and a pre-trained segmentation model consumes them as prompts to generate the final mask (Jiang et al., 23 Aug 2025).

The benchmark itself contains 1,071 reasoning-segmentation examples. Its source imagery comes from six out-of-domain aerial or spaceborne collections: CVUSA, NWPU-RESISC45, CrowdAI, fMoW, CVACT, and LoveDA (Jiang et al., 23 Aug 2025). The initial pool comprises 2,000 images sampled per source, for 12,000 total, after which a quality filter retains images satisfying 40×105\varnothing 40\times 1054, written as

40×105\varnothing 40\times 1055

After automated prompting and human curation, the final size is 1,071 (Jiang et al., 23 Aug 2025).

Each sample contains one natural-language question requiring multi-step geospatial reasoning. The examples given include directional queries such as “Which road lies immediately north of the river bend?”, relational or comparative queries such as “Identify the largest cluster of shipping containers.”, and contextual queries such as “Find the building that is shaded by the tallest tree.” All prompts are generated by Gemini-2.5-Pro and vetted by human annotators to ensure logical complexity (Jiang et al., 23 Aug 2025).

Every example includes a chain-of-thought wrapped in explicit tags: > ... </think> encloses stepwise natural-language reasoning, and <answer> ... </answer> encloses the final grounding in spatial terms. Inside <answer>, the required output schema is a bounding box

40×105\varnothing 40\times 1056

together with two points, 40×105\varnothing 40\times 1057 and 40×105\varnothing 40\times 1058 (Jiang et al., 23 Aug 2025). The paper states that this strict schema enables automatic parsing during RL training.

The segmentation annotations are fine-grained masks 40×105\varnothing 40\times 1059, binary or polygonal, produced by SAM2 or manually in LabelMe. The ground-truth bounding box θ={0 only}\theta=\{0^\circ\text{ only}\}0 is extracted from θ={0 only}\theta=\{0^\circ\text{ only}\}1 as

θ={0 only}\theta=\{0^\circ\text{ only}\}2

with θ={0 only}\theta=\{0^\circ\text{ only}\}3 and θ={0 only}\theta=\{0^\circ\text{ only}\}4. Two positive points are defined: θ={0 only}\theta=\{0^\circ\text{ only}\}5, the center of the maximal inscribed circle in θ={0 only}\theta=\{0^\circ\text{ only}\}6, and θ={0 only}\theta=\{0^\circ\text{ only}\}7, a supplementary point on the mask boundary, described as an outer-ring or farthest-boundary pixel (Jiang et al., 23 Aug 2025).

The split definition is explicit. Training and in-domain test sets are EarthReason, using Million-AID and DIOR, and GeoPixInstruct, using HRSC2016, DOTA-V2.0, and FAIR1M-2.0. The OOD test set is GRASP-1k itself, and no part of GRASP-1k is used during RL training; it serves purely for robustness evaluation (Jiang et al., 23 Aug 2025).

6. Reinforcement-learning signals and benchmark role in the geospatial setting

The core idea of GRASP is to learn bounding-box-plus-point grounding with reinforcement learning from weak spatial cues instead of full masks. For each training question θ={0 only}\theta=\{0^\circ\text{ only}\}8 and image θ={0 only}\theta=\{0^\circ\text{ only}\}9, the multimodal LLM outputs 40\varnothing 400, and rewards are computed against 40\varnothing 401 (Jiang et al., 23 Aug 2025).

Format rewards, denoted 40\varnothing 402, assign 1 point if a well-formed <think>... chain is present and 1 point if the <answer> block exactly follows the bounding-box/point schema (Jiang et al., 23 Aug 2025). Accuracy rewards, denoted 40\varnothing 403, measure grounding quality via bounding-box IoU, a soft scale-normalized box-distance term, and point accuracy. The paper specifies

40\varnothing 404

with reward

40\varnothing 405

The normalized center distance is

40\varnothing 406

and the corresponding reward is

40\varnothing 407

For the points,

40\varnothing 408

with

40\varnothing 409

These rewards drive Grouped Relative Policy Optimization (GRPO) (Jiang et al., 23 Aug 2025).

The paper reports that evaluations on both in-domain and out-of-domain test sets show state-of-the-art results, with about 4% improvement in-domain and up to 54% on OOD benchmarks (Jiang et al., 23 Aug 2025). Within the scope of GRASP-1k specifically, the benchmark is positioned as a robustness test set rather than a training corpus, and the detailed reasoning traces, structured grounding outputs, and fine-grained masks make it suitable for evaluating OOD generalization in geospatial pixel reasoning (Jiang et al., 23 Aug 2025).

7. Comparative interpretation and research significance

The two GRASP-1k datasets occupy very different methodological niches. The robotic GRASP-1k is a physical-interaction dataset built around repeatable object placement, controllable 1-D orientation, and tightly instrumented trial logging (DuFrene et al., 2024). The geospatial GRASP-1k is a reasoning-intensive OOD benchmark organized around language-conditioned grounding, structured intermediate outputs, and reinforcement-learning-compatible reward definitions (Jiang et al., 23 Aug 2025).

Despite the domain difference, both resources formalize an intermediate representation that constrains the task. In robotics, the grasp is parameterized as $105$0 with controlled variation primarily in $105$1 and perturbation $105$2 (DuFrene et al., 2024). In geospatial reasoning, the model must emit a bounded schema consisting of a box and two points before mask generation (Jiang et al., 23 Aug 2025). This suggests a shared design pattern: rather than treating the final task outcome as monolithic, both benchmarks expose structured variables that support fine-grained evaluation.

They also differ in what “benchmarking” means. For the robotic dataset, benchmarking refers to comparing grasp planners on the same 1,020-trial protocol and quantifying success boundaries, repeatability, and success rates under controlled perturbations (DuFrene et al., 2024). For the geospatial dataset, benchmarking refers to robustness evaluation on an OOD set that is excluded from RL training, with performance judged through spatial grounding quality and downstream segmentation behavior (Jiang et al., 23 Aug 2025).

A common misconception would be to treat GRASP-1k as a single canonical benchmark. The literature does not support that reading. Instead, GRASP-1k refers to at least two separate datasets: one for automated robotic grasp trials and one for geospatial pixel reasoning. Accurate citation therefore requires pairing the name with its paper identifier, either (DuFrene et al., 2024) or (Jiang et al., 23 Aug 2025), and with the surrounding task context.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to GRASP-1k.