---
title: Australian Supermarket Object Set (ASOS)
url: https://www.emergentmind.com/topics/australian-supermarket-object-set-asos
type: topic
---

# Australian Supermarket Object Set (ASOS)

Searching arXiv for ASOS and closely related supermarket-object datasets to ground the article with current papers.
The Australian Supermarket Object Set (ASOS) is a dataset of **50 physical supermarket items**, each paired with a **3D watertight textured mesh**, introduced as a practical benchmark for robotics and computer vision [2509.09720]. It is positioned as a response to a recurrent limitation in object-centric benchmarking: many widely used resources provide only digital models, rely on synthetic or specialized objects, or use object sets that are difficult or expensive to obtain in practice. ASOS instead centers on common household items available from a major Australian supermarket chain, Coles, so that researchers can both use the digital assets and acquire the corresponding real objects for physical experiments [2509.09720].

## 1. Dataset identity and design goals

ASOS is explicitly intended to support **reproducible benchmarking across simulation and the real world** while preserving object properties that are hard to simulate faithfully, including **deformability, surface behavior, durability, and mass distribution** [2509.09720]. Its stated application scope includes **object detection, pose estimation, robotic manipulation, and related computer vision and robotics tasks**, with particular relevance to **sim-to-real studies** and experiments in which access to the real objects matters [2509.09720].

A central design principle is accessibility. The dataset focuses on **common household items available from a major Australian supermarket chain, Coles**, rather than synthetic CAD assets or specialized benchmark objects [2509.09720]. The authors argue that this matters because manipulation performance depends not only on geometry and appearance but also on physical properties that are usually absent or poorly approximated in simulation. They also position ASOS relative to established object sets such as **YCB**, noting that standard sets remain valuable but can be harder to obtain internationally as a complete set, and that there is a gap in object collections specifically reflecting supermarket items and varying deformability and mass distributions [2509.09720].

| Property | ASOS characteristic |
|---|---|
| Objects | 50 physical supermarket items |
| Categories | 10 categories |
| Core asset type | 3D watertight textured meshes |
| Object source | Items purchasable from Coles |
| File format | 50 `.ply` files |
| Total construction images | 2,500 images |
| Uncompressed size | 14.6 GB |

This framing makes ASOS less a conventional image-classification benchmark than a physically grounded object set for hybrid evaluation across perception, simulation, and manipulation. A plausible implication is that ASOS is most useful where benchmark fidelity depends on the interplay between object geometry and real-world physical embodiment rather than on image-only recognition performance.

## 2. Object inventory and category structure

The dataset contains **50 objects organized into 10 categories**, with **each category containing between 4 and 6 items** [2509.09720]. The categories are chosen to span household supermarket goods while covering a range of **shapes, scales, and weights that affect graspability** [2509.09720]. The 10 categories are:

- **small boxes**
- **small boxes (food)**
- **large boxes**
- **large boxes (food)**
- **regular cylinders**
- **regular cylinders (food)**
- **irregular cylinders**
- **irregular cylinders (food)**
- **large objects**
- **packets** [2509.09720]

The paper enumerates representative objects within each category. **Small boxes** include **Alcohol Wipes, Ibuprofen, Medistrips, and Tampons**. **Small boxes (food)** include **Electrolyte Tablets, Gravy Mix, Jelly, and Soup Box**. **Large boxes** include **Aluminium Foil, Disposable Gloves, Resealable Bags, Soap Box, and Toothpaste**. **Large boxes (food)** include **Cookies, Crisp Bread, Milk, Rice Bars, Tea, and Water Crackers**. **Regular cylinders** include **Bubbles, Chest Rub Ointment, Disinfectant Bleach, Fish Flakes, Glowsticks, and Rubbish Bags**. **Regular cylinders (food)** include **Baked Beans, Condensed Milk, Diced Tomatoes, Stacked Chips, and Tuna**. **Irregular cylinders** include **Gel Nail Polish Remover, Nail Polish Remover, Sunscreen Roll On, Sunscreen Tube, and Toothbrush Case**. **Irregular cylinders (food)** include **BBQ Sauce, Burger Sauce, Hazelnut Spread, Noodle Cup, and Salt**. **Large objects** include **Bathroom Cleaner, Dishwasher Powder, Dishwashing Liquid, Intimate Wash, and Toilet Cleaner**. **Packets** include **Crackers, Digestive Biscuits, Milk Biscuits, Scotch Finger Biscuits, and Wafers** [2509.09720].

The selection criteria are stated explicitly as **cost, commonality, shape/size/weight diversity, and variety of use** [2509.09720]. Cost is addressed by restricting the set to Coles-purchasable items, favoring **cheap generic brands** and **non-perishable, robust items**. Commonality is addressed by selecting **commonly purchased generic goods**. Shape, size, and weight diversity are emphasized because they directly affect grasping and manipulation. Variety is introduced by covering **food items, drinks, cleaning goods, personal hygiene products, and health items**, while also including outliers such as **deformable biscuit packets** and **irregular spray-bottle-like containers** [2509.09720].

The category design is strongly morphology-oriented. The explicit subdivision of boxes and cylinders into smaller structural classes, and the separate inclusion of packets and large objects, suggests that ASOS is intended to expose grasping and perception systems to packaging geometries that differ not merely semantically but operationally. This suggests a benchmark philosophy organized around manipulation-relevant object form rather than around retail ontology alone.

## 3. Physical metadata and object properties

Each object is accompanied by **detailed information about its mass and dimensions** [2509.09720]. Across the full set, **mass ranges from 13 g for Medistrips to 1458 g for Disinfectant Bleach Cleaner**; the paper also summarizes the textual range as from **18 g to 1458 g**, but the tabulated value for Medistrips is the most specific figure reported [2509.09720]. The upper end is described as deliberately kept below the maximum payload of most standard robotic manipulators [2509.09720].

Dimensions are reported in **millimeters**. For boxes the representation is **width × height × depth**, and for cylinders it is **height × diameter**, as specified in the table caption [2509.09720]. Examples reported in the paper include:

- **Medistrips** at \(74 \times 96 \times 22\) mm
- **Aluminium Foil** at \(312 \times 51 \times 52\) mm
- **Disinfectant Bleach** at \(293 \times 85\) mm
- **Milk** at \(92 \times 197 \times 58\) mm and **1074 g**
- **Dishwasher Powder** at \(140 \times 220 \times 65\) mm and **1128 g**
- **Stacked Chips** at \(233 \times 72\) mm and **183 g**
- **Toothbrush Case** at \(30 \times 208 \times 20\) mm and **24 g** [2509.09720]

The dataset does **not** report quantified aggregate statistics for **material classes, packaging materials, reflectiveness, texture entropy, rigidity/deformability labels, or shape complexity metrics** [2509.09720]. It also does **not** provide formal annotations for **“rigid” versus “deformable”**, nor does it provide **inertial tensor** or **center-of-mass** data [2509.09720]. The paper nevertheless highlights qualitative diversity along several of these axes, explicitly mentioning **varying deformability and mass distribution** as motivations for the set and describing the inclusion of **deformable packet-like objects** and **irregularly shaped cleaning or sauce containers** [2509.09720].

This combination of measured mass/size metadata and absent richer physical labels is significant. ASOS provides enough information for basic physical characterization and simulator scaling, but not the full physical parameterization often required for high-fidelity dynamics modeling. A plausible implication is that the dataset is immediately useful for geometry-centric benchmarking, while more exact physically based simulation still requires user-supplied estimates or calibration.

## 4. 3D acquisition and reconstruction pipeline

The 3D acquisition workflow is built around a **Structure-from-Motion and multi-view stereo approach based on COLMAP**, citing both the COLMAP SfM and MVS papers [2509.09720]. For each object, the object is placed inside a **“feature-rich box”** containing various colored shapes that create a background rich in visual features; the authors state that this simplifies feature detection and matching [2509.09720]. The paper explicitly mentions **SIFT** for feature extraction and matching and **RANSAC** for robust correspondence and model estimation [2509.09720].

The capture hardware is an **iPhone 13 mini** [2509.09720]. The capture protocol is viewpoint-based rather than turntable-based. The **front and side** of each object are photographed from **25 different views spanning a semicircle** around the object, with **one photo per view**, yielding:

$$50 \text{ photos per object} = 25 \text{ views} \times 2 \text{ sides}.$$

The image resolution is:

$$4032 \times 3024.$$

The paper does **not** specify focal length settings, exposure settings, lighting hardware, or a dedicated camera calibration process beyond COLMAP-based camera pose recovery [2509.09720]. It also does **not** mention segmentation during capture, chroma-key backgrounds, fiducial markers, or controlled studio illumination [2509.09720].

After image capture, **COLMAP** reconstructs a high-quality point cloud of the scene. The object is then **isolated from the rest of the scene and cleaned using MeshLab tools**. Once an isolated object point cloud is obtained, **screened Poisson surface reconstruction** is applied to create a **watertight mesh**, and the resulting mesh is subsequently cleaned because Poisson reconstruction can introduce artifacts [2509.09720]. Since the bottom of the object is not visible when resting on the table, the acquisition is repeated after **flipping the object** to expose previously unseen surfaces. The two partial meshes are then merged using **point-based gluing via the Iterative Closest Point (ICP) algorithm**, producing the final watertight object mesh [2509.09720].

The paper summarizes the pipeline as: **image capture; SfM camera recovery and sparse reconstruction; dense point cloud generation; object isolation and cleanup; screened Poisson reconstruction into a watertight surface; artifact cleanup; second-pass scanning after flipping the object; ICP-based registration and gluing of the two halves; and final watertight mesh generation** [2509.09720].

At the same time, the paper does **not** provide quantitative reconstruction parameters such as **COLMAP feature thresholds, bundle adjustment settings, dense stereo parameters, Poisson octree depth, screening weight, ICP convergence thresholds, mesh decimation targets, or alignment tolerances** [2509.09720]. It also does **not** detail **UV unwrapping** or **texture baking settings**, despite repeatedly referring to the models as **“textured meshes”** [2509.09720]. This makes the pipeline reproducible at the algorithmic level but not fully specified at the implementation-parameter level.

## 5. Released assets, annotation scope, and benchmark affordances

The final released object set consists of **50 `.ply` files**, one per object, together with metadata about **mass and dimensions** [2509.09720]. The full dataset was built from **2,500 images total**, corresponding to **50 images per object**, and occupies **14.6 GB uncompressed** [2509.09720]. The paper states that the object set data is accessible online and provides a public project webpage:

\[
\texttt{https://lachlanchumbley.github.io/ColesObjectSet/}
\]

The annotation scope is intentionally narrow. The paper does **not** state that the raw images are released, although it reports their number and resolution [2509.09720]. It also does **not** describe the release of **camera poses, object poses, segmentation masks, depth maps, bounding boxes, object symmetries, inertial properties, train/validation/test splits, or benchmark protocol files** [2509.09720]. No **coordinate-frame convention, object origin definition, axis convention, scale convention, or naming schema beyond the human-readable object names** is specified [2509.09720]. Units are implicit in the object property table, where mass is in **grams** and dimensions are in **millimeters** [2509.09720].

The paper further notes that there is **no explicit statement** about whether the meshes are metrically scaled in the `.ply` files, although the existence of measured dimensions strongly suggests that physical scaling is part of the metadata; still, the paper does not formally define mesh units or origin [2509.09720]. Likewise, ASOS does **not** include simulator-ready artifacts such as **URDFs, collision meshes, articulated models, or precomputed inertial parameters** [2509.09720].

These omissions matter for benchmark interpretation. ASOS is presented primarily as an **object set** rather than as a completed challenge suite. Notably, the paper does **not** include a formal benchmark suite with task-specific metrics, and it reports **no baseline experiments** for **detection, 6D pose estimation, grasp planning, or manipulation success** [2509.09720]. The contribution is therefore infrastructural and methodological rather than leaderboard-oriented.

## 6. Position within supermarket and robotics dataset research

ASOS is framed as complementary to object and retail datasets such as **YCB, ACRV, LINEMOD, BigBIRD, GSO, and ABO** [2509.09720]. Its distinctive contribution is not richer visual annotation than those resources, but rather the coupling of **purchasable real objects** with **watertight digital meshes** in a supermarket-object domain [2509.09720]. Compared with online-only repositories, this is presented as particularly important for manipulation tasks where **deformability, frictional surface effects, durability, and mass-related behavior** influence outcomes [2509.09720].

Within supermarket-specific vision research, other datasets emphasize different problem formulations. **MVTec D2S** is a benchmark for **instance-aware semantic segmentation** with **21,000 high-resolution images**, **60 object categories**, and **pixel-wise labels of all object instances**, designed for **automatic checkout, inventory, or warehouse systems** [1804.08292]. Its split deliberately creates domain shift between simple training scenes and more complex validation/test scenes [1804.08292]. By contrast, ASOS does not provide scene imagery, instance masks, or a train/validation/test protocol, and therefore occupies a different point in the design space: it is an object-centered physical benchmark rather than a dense scene-understanding dataset.

Likewise, **“Acquire, Augment, Segment & Enjoy: Weakly Supervised Instance Segmentation of Supermarket Products”** focuses on methodology for generating training annotations from controlled acquisition, using **D2S** and a **Mask R-CNN / Detectron** pipeline rather than ASOS itself [1807.02001]. That work is relevant because it demonstrates how supermarket-product datasets can be built around **controlled isolated capture + automatic mask extraction + synthetic clutter generation** [1807.02001]. ASOS does not itself implement such a segmentation benchmark, but a plausible implication is that its physical objects and meshes could support analogous downstream data-generation workflows if additional scene-level imaging and labels were created.

A particularly close downstream connection appears in **Supermarket-6DoF**, which introduces **1,500 real-world grasp attempts across 20 supermarket objects with publicly available 3D models** and states that the 20 objects were selected from a **Supermarket Object Set** [2502.16311]. That paper does **not explicitly mention ASOS by name**, nor does it claim terminological equivalence, but it is highly relevant to ASOS-style robotic manipulation benchmarks [2502.16311]. It adds **single-view eye-in-hand RGB**, **depth**, **point clouds**, **full 6-DoF grasp poses**, and binary labels for **grasp success** and **grasp robustness/stability** [2502.16311]. This suggests one route by which an object set like ASOS can evolve into an execution-grounded manipulation dataset.

## 7. Limitations, caveats, and prospective extensions

The limitations of ASOS are explicitly stated. Although the digital meshes enable simulation, some physical properties remain difficult to simulate effectively, especially **flexibility, deformability, durability, and mass distribution** [2509.09720]. The authors note that **mass distribution is difficult to measure accurately**, and no such measurements are provided [2509.09720]. Another practical limitation is that the scan pipeline requires **flipping objects** because the support surface occludes the bottom face, so final meshes depend on registration of two separate scans; the paper does **not** quantify any resulting reconstruction errors [2509.09720].

The annotation scope is also limited. Beyond mesh geometry and basic physical dimensions and masses, the dataset provides no rich supervision for **pose estimation**, **detection benchmarking**, or **scene understanding** unless users create additional labeled imagery themselves [2509.09720]. The paper does **not** specify **licensing terms** for the digital files, **versioning**, **maintenance plans**, or **update schedules** [2509.09720]. It also does not state a total replacement cost for the full set [2509.09720].

The paper nonetheless indicates future expansion directions. Its conclusion states that future work could extend the dataset with **additional categories** or with **domain-specific tasks such as grocery sorting or packaging automation** [2509.09720]. This suggests that ASOS, in its current form, should be understood as a foundational object collection rather than a finished end-to-end benchmark ecosystem.

In that sense, ASOS occupies a specific niche. It provides **real, purchasable Australian supermarket items**, **watertight textured 3D meshes**, and **measured mass and dimension metadata** [2509.09720]. It does not yet provide the richer annotations, benchmark splits, or formal evaluation protocols typical of mature task-centric datasets. Its main significance lies in creating an accessible and reproducible bridge between **simulation assets** and **physical supermarket objects** for robotics and computer vision research [2509.09720].

Source: https://www.emergentmind.com/topics/australian-supermarket-object-set-asos