Papers
Topics
Authors
Recent
Search
2000 character limit reached

Australian Supermarket Object Set (ASOS)

Updated 10 July 2026
  • ASOS is a dataset of 50 common supermarket items with watertight textured 3D meshes designed to benchmark sim-to-real robotics and computer vision tasks.
  • It employs a COLMAP-based acquisition pipeline with ICP registration to merge multi-view scans for accurate object reconstruction.
  • The set bridges simulation and physical experiments by providing measured mass and dimensions, essential for evaluating graspability and manipulation.

Searching arXiv for ASOS and closely related supermarket-object datasets to ground the article with current papers. The Australian Supermarket Object Set (ASOS) is a dataset of 50 physical supermarket items, each paired with a 3D watertight textured mesh, introduced as a practical benchmark for robotics and computer vision (Cosgun et al., 9 Sep 2025). It is positioned as a response to a recurrent limitation in object-centric benchmarking: many widely used resources provide only digital models, rely on synthetic or specialized objects, or use object sets that are difficult or expensive to obtain in practice. ASOS instead centers on common household items available from a major Australian supermarket chain, Coles, so that researchers can both use the digital assets and acquire the corresponding real objects for physical experiments (Cosgun et al., 9 Sep 2025).

1. Dataset identity and design goals

ASOS is explicitly intended to support reproducible benchmarking across simulation and the real world while preserving object properties that are hard to simulate faithfully, including deformability, surface behavior, durability, and mass distribution (Cosgun et al., 9 Sep 2025). Its stated application scope includes object detection, pose estimation, robotic manipulation, and related computer vision and robotics tasks, with particular relevance to sim-to-real studies and experiments in which access to the real objects matters (Cosgun et al., 9 Sep 2025).

A central design principle is accessibility. The dataset focuses on common household items available from a major Australian supermarket chain, Coles, rather than synthetic CAD assets or specialized benchmark objects (Cosgun et al., 9 Sep 2025). The authors argue that this matters because manipulation performance depends not only on geometry and appearance but also on physical properties that are usually absent or poorly approximated in simulation. They also position ASOS relative to established object sets such as YCB, noting that standard sets remain valuable but can be harder to obtain internationally as a complete set, and that there is a gap in object collections specifically reflecting supermarket items and varying deformability and mass distributions (Cosgun et al., 9 Sep 2025).

Property ASOS characteristic
Objects 50 physical supermarket items
Categories 10 categories
Core asset type 3D watertight textured meshes
Object source Items purchasable from Coles
File format 50 .ply files
Total construction images 2,500 images
Uncompressed size 14.6 GB

This framing makes ASOS less a conventional image-classification benchmark than a physically grounded object set for hybrid evaluation across perception, simulation, and manipulation. A plausible implication is that ASOS is most useful where benchmark fidelity depends on the interplay between object geometry and real-world physical embodiment rather than on image-only recognition performance.

2. Object inventory and category structure

The dataset contains 50 objects organized into 10 categories, with each category containing between 4 and 6 items (Cosgun et al., 9 Sep 2025). The categories are chosen to span household supermarket goods while covering a range of shapes, scales, and weights that affect graspability (Cosgun et al., 9 Sep 2025). The 10 categories are:

  • small boxes
  • small boxes (food)
  • large boxes
  • large boxes (food)
  • regular cylinders
  • regular cylinders (food)
  • irregular cylinders
  • irregular cylinders (food)
  • large objects
  • packets (Cosgun et al., 9 Sep 2025)

The paper enumerates representative objects within each category. Small boxes include Alcohol Wipes, Ibuprofen, Medistrips, and Tampons. Small boxes (food) include Electrolyte Tablets, Gravy Mix, Jelly, and Soup Box. Large boxes include Aluminium Foil, Disposable Gloves, Resealable Bags, Soap Box, and Toothpaste. Large boxes (food) include Cookies, Crisp Bread, Milk, Rice Bars, Tea, and Water Crackers. Regular cylinders include Bubbles, Chest Rub Ointment, Disinfectant Bleach, Fish Flakes, Glowsticks, and Rubbish Bags. Regular cylinders (food) include Baked Beans, Condensed Milk, Diced Tomatoes, Stacked Chips, and Tuna. Irregular cylinders include Gel Nail Polish Remover, Nail Polish Remover, Sunscreen Roll On, Sunscreen Tube, and Toothbrush Case. Irregular cylinders (food) include BBQ Sauce, Burger Sauce, Hazelnut Spread, Noodle Cup, and Salt. Large objects include Bathroom Cleaner, Dishwasher Powder, Dishwashing Liquid, Intimate Wash, and Toilet Cleaner. Packets include Crackers, Digestive Biscuits, Milk Biscuits, Scotch Finger Biscuits, and Wafers (Cosgun et al., 9 Sep 2025).

The selection criteria are stated explicitly as cost, commonality, shape/size/weight diversity, and variety of use (Cosgun et al., 9 Sep 2025). Cost is addressed by restricting the set to Coles-purchasable items, favoring cheap generic brands and non-perishable, robust items. Commonality is addressed by selecting commonly purchased generic goods. Shape, size, and weight diversity are emphasized because they directly affect grasping and manipulation. Variety is introduced by covering food items, drinks, cleaning goods, personal hygiene products, and health items, while also including outliers such as deformable biscuit packets and irregular spray-bottle-like containers (Cosgun et al., 9 Sep 2025).

The category design is strongly morphology-oriented. The explicit subdivision of boxes and cylinders into smaller structural classes, and the separate inclusion of packets and large objects, suggests that ASOS is intended to expose grasping and perception systems to packaging geometries that differ not merely semantically but operationally. This suggests a benchmark philosophy organized around manipulation-relevant object form rather than around retail ontology alone.

3. Physical metadata and object properties

Each object is accompanied by detailed information about its mass and dimensions (Cosgun et al., 9 Sep 2025). Across the full set, mass ranges from 13 g for Medistrips to 1458 g for Disinfectant Bleach Cleaner; the paper also summarizes the textual range as from 18 g to 1458 g, but the tabulated value for Medistrips is the most specific figure reported (Cosgun et al., 9 Sep 2025). The upper end is described as deliberately kept below the maximum payload of most standard robotic manipulators (Cosgun et al., 9 Sep 2025).

Dimensions are reported in millimeters. For boxes the representation is width × height × depth, and for cylinders it is height × diameter, as specified in the table caption (Cosgun et al., 9 Sep 2025). Examples reported in the paper include:

  • Medistrips at 74×96×2274 \times 96 \times 22 mm
  • Aluminium Foil at 312×51×52312 \times 51 \times 52 mm
  • Disinfectant Bleach at 293×85293 \times 85 mm
  • Milk at 92×197×5892 \times 197 \times 58 mm and 1074 g
  • Dishwasher Powder at 140×220×65140 \times 220 \times 65 mm and 1128 g
  • Stacked Chips at 233×72233 \times 72 mm and 183 g
  • Toothbrush Case at 30×208×2030 \times 208 \times 20 mm and 24 g (Cosgun et al., 9 Sep 2025)

The dataset does not report quantified aggregate statistics for material classes, packaging materials, reflectiveness, texture entropy, rigidity/deformability labels, or shape complexity metrics (Cosgun et al., 9 Sep 2025). It also does not provide formal annotations for “rigid” versus “deformable”, nor does it provide inertial tensor or center-of-mass data (Cosgun et al., 9 Sep 2025). The paper nevertheless highlights qualitative diversity along several of these axes, explicitly mentioning varying deformability and mass distribution as motivations for the set and describing the inclusion of deformable packet-like objects and irregularly shaped cleaning or sauce containers (Cosgun et al., 9 Sep 2025).

This combination of measured mass/size metadata and absent richer physical labels is significant. ASOS provides enough information for basic physical characterization and simulator scaling, but not the full physical parameterization often required for high-fidelity dynamics modeling. A plausible implication is that the dataset is immediately useful for geometry-centric benchmarking, while more exact physically based simulation still requires user-supplied estimates or calibration.

4. 3D acquisition and reconstruction pipeline

The 3D acquisition workflow is built around a Structure-from-Motion and multi-view stereo approach based on COLMAP, citing both the COLMAP SfM and MVS papers (Cosgun et al., 9 Sep 2025). For each object, the object is placed inside a “feature-rich box” containing various colored shapes that create a background rich in visual features; the authors state that this simplifies feature detection and matching (Cosgun et al., 9 Sep 2025). The paper explicitly mentions SIFT for feature extraction and matching and RANSAC for robust correspondence and model estimation (Cosgun et al., 9 Sep 2025).

The capture hardware is an iPhone 13 mini (Cosgun et al., 9 Sep 2025). The capture protocol is viewpoint-based rather than turntable-based. The front and side of each object are photographed from 25 different views spanning a semicircle around the object, with one photo per view, yielding:

50 photos per object=25 views×2 sides.50 \text{ photos per object} = 25 \text{ views} \times 2 \text{ sides}.

The image resolution is:

4032×3024.4032 \times 3024.

The paper does not specify focal length settings, exposure settings, lighting hardware, or a dedicated camera calibration process beyond COLMAP-based camera pose recovery (Cosgun et al., 9 Sep 2025). It also does not mention segmentation during capture, chroma-key backgrounds, fiducial markers, or controlled studio illumination (Cosgun et al., 9 Sep 2025).

After image capture, COLMAP reconstructs a high-quality point cloud of the scene. The object is then isolated from the rest of the scene and cleaned using MeshLab tools. Once an isolated object point cloud is obtained, screened Poisson surface reconstruction is applied to create a watertight mesh, and the resulting mesh is subsequently cleaned because Poisson reconstruction can introduce artifacts (Cosgun et al., 9 Sep 2025). Since the bottom of the object is not visible when resting on the table, the acquisition is repeated after flipping the object to expose previously unseen surfaces. The two partial meshes are then merged using point-based gluing via the Iterative Closest Point (ICP) algorithm, producing the final watertight object mesh (Cosgun et al., 9 Sep 2025).

The paper summarizes the pipeline as: image capture; SfM camera recovery and sparse reconstruction; dense point cloud generation; object isolation and cleanup; screened Poisson reconstruction into a watertight surface; artifact cleanup; second-pass scanning after flipping the object; ICP-based registration and gluing of the two halves; and final watertight mesh generation (Cosgun et al., 9 Sep 2025).

At the same time, the paper does not provide quantitative reconstruction parameters such as COLMAP feature thresholds, bundle adjustment settings, dense stereo parameters, Poisson octree depth, screening weight, ICP convergence thresholds, mesh decimation targets, or alignment tolerances (Cosgun et al., 9 Sep 2025). It also does not detail UV unwrapping or texture baking settings, despite repeatedly referring to the models as “textured meshes” (Cosgun et al., 9 Sep 2025). This makes the pipeline reproducible at the algorithmic level but not fully specified at the implementation-parameter level.

5. Released assets, annotation scope, and benchmark affordances

The final released object set consists of 50 .ply files, one per object, together with metadata about mass and dimensions (Cosgun et al., 9 Sep 2025). The full dataset was built from 2,500 images total, corresponding to 50 images per object, and occupies 14.6 GB uncompressed (Cosgun et al., 9 Sep 2025). The paper states that the object set data is accessible online and provides a public project webpage:

https://lachlanchumbley.github.io/ColesObjectSet/\texttt{https://lachlanchumbley.github.io/ColesObjectSet/}

The annotation scope is intentionally narrow. The paper does not state that the raw images are released, although it reports their number and resolution (Cosgun et al., 9 Sep 2025). It also does not describe the release of camera poses, object poses, segmentation masks, depth maps, bounding boxes, object symmetries, inertial properties, train/validation/test splits, or benchmark protocol files (Cosgun et al., 9 Sep 2025). No coordinate-frame convention, object origin definition, axis convention, scale convention, or naming schema beyond the human-readable object names is specified (Cosgun et al., 9 Sep 2025). Units are implicit in the object property table, where mass is in grams and dimensions are in millimeters (Cosgun et al., 9 Sep 2025).

The paper further notes that there is no explicit statement about whether the meshes are metrically scaled in the .ply files, although the existence of measured dimensions strongly suggests that physical scaling is part of the metadata; still, the paper does not formally define mesh units or origin (Cosgun et al., 9 Sep 2025). Likewise, ASOS does not include simulator-ready artifacts such as URDFs, collision meshes, articulated models, or precomputed inertial parameters (Cosgun et al., 9 Sep 2025).

These omissions matter for benchmark interpretation. ASOS is presented primarily as an object set rather than as a completed challenge suite. Notably, the paper does not include a formal benchmark suite with task-specific metrics, and it reports no baseline experiments for detection, 6D pose estimation, grasp planning, or manipulation success (Cosgun et al., 9 Sep 2025). The contribution is therefore infrastructural and methodological rather than leaderboard-oriented.

6. Position within supermarket and robotics dataset research

ASOS is framed as complementary to object and retail datasets such as YCB, ACRV, LINEMOD, BigBIRD, GSO, and ABO (Cosgun et al., 9 Sep 2025). Its distinctive contribution is not richer visual annotation than those resources, but rather the coupling of purchasable real objects with watertight digital meshes in a supermarket-object domain (Cosgun et al., 9 Sep 2025). Compared with online-only repositories, this is presented as particularly important for manipulation tasks where deformability, frictional surface effects, durability, and mass-related behavior influence outcomes (Cosgun et al., 9 Sep 2025).

Within supermarket-specific vision research, other datasets emphasize different problem formulations. MVTec D2S is a benchmark for instance-aware semantic segmentation with 21,000 high-resolution images, 60 object categories, and pixel-wise labels of all object instances, designed for automatic checkout, inventory, or warehouse systems (Follmann et al., 2018). Its split deliberately creates domain shift between simple training scenes and more complex validation/test scenes (Follmann et al., 2018). By contrast, ASOS does not provide scene imagery, instance masks, or a train/validation/test protocol, and therefore occupies a different point in the design space: it is an object-centered physical benchmark rather than a dense scene-understanding dataset.

Likewise, “Acquire, Augment, Segment & Enjoy: Weakly Supervised Instance Segmentation of Supermarket Products” focuses on methodology for generating training annotations from controlled acquisition, using D2S and a Mask R-CNN / Detectron pipeline rather than ASOS itself (Follmann et al., 2018). That work is relevant because it demonstrates how supermarket-product datasets can be built around controlled isolated capture + automatic mask extraction + synthetic clutter generation (Follmann et al., 2018). ASOS does not itself implement such a segmentation benchmark, but a plausible implication is that its physical objects and meshes could support analogous downstream data-generation workflows if additional scene-level imaging and labels were created.

A particularly close downstream connection appears in Supermarket-6DoF, which introduces 1,500 real-world grasp attempts across 20 supermarket objects with publicly available 3D models and states that the 20 objects were selected from a Supermarket Object Set (Toskov et al., 22 Feb 2025). That paper does not explicitly mention ASOS by name, nor does it claim terminological equivalence, but it is highly relevant to ASOS-style robotic manipulation benchmarks (Toskov et al., 22 Feb 2025). It adds single-view eye-in-hand RGB, depth, point clouds, full 6-DoF grasp poses, and binary labels for grasp success and grasp robustness/stability (Toskov et al., 22 Feb 2025). This suggests one route by which an object set like ASOS can evolve into an execution-grounded manipulation dataset.

7. Limitations, caveats, and prospective extensions

The limitations of ASOS are explicitly stated. Although the digital meshes enable simulation, some physical properties remain difficult to simulate effectively, especially flexibility, deformability, durability, and mass distribution (Cosgun et al., 9 Sep 2025). The authors note that mass distribution is difficult to measure accurately, and no such measurements are provided (Cosgun et al., 9 Sep 2025). Another practical limitation is that the scan pipeline requires flipping objects because the support surface occludes the bottom face, so final meshes depend on registration of two separate scans; the paper does not quantify any resulting reconstruction errors (Cosgun et al., 9 Sep 2025).

The annotation scope is also limited. Beyond mesh geometry and basic physical dimensions and masses, the dataset provides no rich supervision for pose estimation, detection benchmarking, or scene understanding unless users create additional labeled imagery themselves (Cosgun et al., 9 Sep 2025). The paper does not specify licensing terms for the digital files, versioning, maintenance plans, or update schedules (Cosgun et al., 9 Sep 2025). It also does not state a total replacement cost for the full set (Cosgun et al., 9 Sep 2025).

The paper nonetheless indicates future expansion directions. Its conclusion states that future work could extend the dataset with additional categories or with domain-specific tasks such as grocery sorting or packaging automation (Cosgun et al., 9 Sep 2025). This suggests that ASOS, in its current form, should be understood as a foundational object collection rather than a finished end-to-end benchmark ecosystem.

In that sense, ASOS occupies a specific niche. It provides real, purchasable Australian supermarket items, watertight textured 3D meshes, and measured mass and dimension metadata (Cosgun et al., 9 Sep 2025). It does not yet provide the richer annotations, benchmark splits, or formal evaluation protocols typical of mature task-centric datasets. Its main significance lies in creating an accessible and reproducible bridge between simulation assets and physical supermarket objects for robotics and computer vision research (Cosgun et al., 9 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Australian Supermarket Object Set (ASOS).