---
title: Hierarchical Open-Vocabulary Scene Graphs
url: https://www.emergentmind.com/topics/hierarchical-open-vocabulary-scene-graphs-hov-sg
type: topic
---

# Hierarchical Open-Vocabulary Scene Graphs

A Hierarchical Open-Vocabulary Scene Graph (HOV-SG) is a structured, multi-level graph-based representation designed to model the semantic and relational structure of complex visual environments—spanning both 2D images and 3D spaces—without restricting semantics or relationships to a fixed, closed vocabulary. HOV-SG integrates foundation vision-language models, large language models (LLMs), and hierarchical abstraction, yielding a formalism that enables zero-shot category extension, interpretable spatial structure, efficient scene reasoning, and robust downstream application to navigation, grounding, and scene synthesis. This paradigm has driven significant advances in open-world perception, robotic scene understanding, and language-driven reasoning.

## 1. Formal Structure and Semantics of HOV-SG

At its core, HOV-SG represents a scene as a multi-layer graph $\mathcal G = (V, E, A, L)$ where nodes $V$ are partitioned into abstraction levels, e.g., objects, rooms, floors, and the building (or, in outdoor settings, objects, road segments, intersections, and environment). Edges $E$ include intra-level (e.g., adjacency), inter-level (e.g., parent-child), semantic (e.g., “on top of,” “contains”), and spatial/proximity relations. Each node and edge is associated with feature attributes $A_V,$ $A_E$, typically open-vocabulary CLIP- or VLM-derived embeddings, and an abstraction level label $L$ [2403.17846], [2602.02456], [2502.10675], [2409.10350], [2410.06239], [2403.09412], [2503.08474].

For 3D or embodied scenes, typical node levels are:
- $V^{(1)}$: Object instances with open-vocabulary semantic descriptors
- $V^{(2)}$: Place/room segmentation (often via clustering or geometric partitioning)
- $V^{(3)}$: Floors (or, for outdoor, road segments, lane structures)
- $V^{(4)}$: Root or environment node

This hierarchical construction enforces spatial containment, enables fast retrieval (object $\rightarrow$ room $\rightarrow$ floor), and supports reasoning over both fine and coarse semantic levels.

## 2. Construction Methodologies and Pipeline Components

### 2.1. Hierarchical Scene Graph Construction

Construction proceeds from raw sensor data (RGB(-D), LiDAR, or point clouds) using a series of geometric, topological, and semantic partitioning steps:
- **Semantic Segmentation and Mask Projection**: Advanced segmentation (SAM, Grounding DINO, TAP) produces 2D or 3D masks, which are then back-projected into the global frame using odometry or extrinsics [2403.17846], [2503.08474], [2403.09412].
- **Instance/Node Formation and Semantic Labeling**: Masked regions are encoded with vision-language models (CLIP, BLIP, Sentence-BERT, or Uni3D), yielding open-vocabulary, language-aligned feature vectors for each object or region [2412.19021], [2409.10350].
- **Hierarchical Partitioning**: Floor, room, and segment boundaries are defined using clustering in the height dimension (for floors), watershed on BEV-projected points (for rooms), or geometric/DBSCAN approaches (for outdoor segments or urban lane-graphs) [2403.17846], [2403.09412], [2503.08474].
- **Edge Construction**: Hierarchical parent–child edges, adjacency and spatial relations (e.g., “left-of” via centroid delta), and (if available) semantic relationship edges via vision–language relation heads [2409.10350], [2412.19021].

A representative scene graph instantiation:
| Level       | Node Type            | Edges                         | Feature Attribute        |
|-------------|---------------------|-------------------------------|-------------------------|
| $V^{(1)}$   | Object instances    | intra-object rel, to room     | CLIP/VLM object embedding |
| $V^{(2)}$   | Room/segment        | to contained objects, adj.    | Room-level CLIP cluster   |
| $V^{(3)}$   | Floor/road segment  | to rooms below, to other segs | Floor label embedding     |
| $V^{(4)}$   | Root/environment    | to all floors                 | Environment descriptor    |

### 2.2. Open-Vocabulary Semantic Mapping

Open-vocabulary assignment leverages joint vision–language embedding spaces, generally freezing the CLIP/BLIP/VLM backbone and ranking visual features against arbitrary text queries via cosine similarity. Multi-crop fusion, k-means over view clusters, or Sentence-BERT encodings are common techniques for robust per-node semantics [2403.17846], [2403.09412], [2502.10675].

For relationships, open-vocabulary predicate heads (Bayesian or hierarchical) and entity-aware/region-aware hierarchical prompts (with LLM mining) enable the extension to zero-shot predicates while maintaining semantic structure and strong performance on novel relations [2303.06842], [2412.19021].

### 2.3. Graph Optimization and Update Algorithms

Dynamic scene graphs in multi-agent or dynamic environments employ incremental association, fusion, and relabeling schemes: keyframe pose graph optimization, DBSCAN-based association, IoU or embedding-similarity-based object fusion, and periodic semantic updates [2503.08474], [2410.06239].

Memory efficiency is achieved by storing only per-segment/room CLIP features, leading to a ≈ 75% reduction in storage relative to dense voxel-based alternatives, while retaining global retrieval capabilities [2403.17846], [2403.09412].

## 3. Integration with Vision-Language Models and LLM Reasoning

HOV-SG frameworks exploit foundation models at multiple stages:

- **Vision-Language Models (VLMs)** provide dense, open-vocabulary feature representations at node and relation levels, underpinning zero-shot identification and retrieval [2403.09412], [2403.17846].
- **Hierarchical Prompting and LLM Mining**: Hierarchical relation and entity clustering uses LLMs to mine fine-grained, region-aware prompts. This two-level prompt structure (entity-aware + region-aware) boosts novel predicate alignment and model robustness [2412.19021].
- **Task and Query Reasoning via LLMs**: For downstream query parsing, room-type designation, and multi-step task planning, LLMs process graph signatures or node/edge features, decompose commands (chain-of-thought), and coordinate plan generation in natural language [2602.02456], [2410.06239], [2502.10675].

This tight coupling yields high-level scene understanding, supports spatial referencing across hierarchical layers (room to object, object to object), and enables robust grounding of natural language queries in complex environments [2507.12123], [2602.02456], [2502.10675].

## 4. Quantitative Results and Evaluation Protocols

Experiments consistently demonstrate that HOV-SG methods outperform closed-set or flat-vocabulary baselines in zero-shot segmentation, open-vocabulary retrieval, semantic alignment, and physical feasibility:
- **Object/Room Segmentation and Retrieval**: On HM3DSem, HOV-SG achieves AUC$_k^{top}$ of 84.9% (ConceptGraphs 84.1%, VLMaps 56.2%) [2403.17846].
- **Open-Vocabulary Segmentation**: On SemanticKITTI, OpenGraph mIoU (seq 03) 0.6051, F1 0.7302, outperforming RangeNet++ and DeepLab V3 [2403.09412].
- **Physical Feasibility Metrics**: HOV-SG yields 0% OOB and overlap, KL-divergence 0.09, compared to ATISS (OOB 0.48%) and DiffuScene (OOB 0.77%) [2502.10675].
- **SGG Benchmarks**: On Visual Genome, hierarchical relation heads (Bayesian or prompt-based) increase mR@50 by >6 points versus flat heads and yield a zero-shot PredCLS R@50 of 20.4 (vs baseline 3.6–15.1) [2303.06842], [2412.19021].
- **Navigation Success**: In real-world trials, hierarchical open-vocabulary scene graphs enable robot navigation with 56.1–100% task completion for object, room, and floor-level goals [2403.17846], [2410.06239].

| Setting                          | Key Metric                          | HOV-SG Result         | Baseline         |
|-----------------------------------|-------------------------------------|-----------------------|------------------|
| HM3DSem (obj ret.)                | AUC$_k^{top}$                       | 84.9%                | 84.1% / 56.2%    |
| SemanticKITTI (open-seg)          | mIoU / F1 (seq 03)                  | 0.6051 / 0.7302      | 0.4780 / 0.6115  |
| Scene synthesis (OOB/overlap/KL)  | 0.0% / 0.0% / 0.09                  | 0.48% / 0.18% / 0.19 |
| Visual Genome (PredCLS-zsR@50)    | 20.4                                | 3.6–15.1             |

## 5. Open-Vocabulary Relation Modeling and Hierarchical Prompting

HOV-SG advances open-vocabulary relation modeling by hierarchically structuring both the relation label space (using super-categories, e.g., geometric, possessive, semantic [2303.06842]) and the textual representation pool (super-entities and region-level prompts [2412.19021]):
- **Bayesian Relation Heads**: Factorize relation prediction via $P(r) = P(\text{super}=s) P(r|s)$, supporting seamless insertion of novel predicates and zero-shot retrieval [2303.06842].
- **Hierarchical Prompt Pools**: RAHP constructs two-level prompt libraries: (1) entity-aware: subject–object super-pair + predicate, (2) region-aware: LLM-mined spatially or functionally specific prompts [2412.19021]. Dynamic selection mechanisms efficiently prune noise.
- **Clustering and Prompt Mining**: Entities are clustered via text encoder and k-means, with LLMs generating human-interpretable super-entity names and region descriptions, balancing diversity and computational cost.

RAHP and similar frameworks show 4–8 points mR@100 improvement on Visual Genome and Open Images v6, consistently increasing recall on novel predicate types.

## 6. Real-World Applications and System-Level Integration

HOV-SG underpins robust scene understanding pipelines in diverse embodied and virtual domains:
- **Language-Conditioned Navigation**: HOV-SG supplies hierarchical grounding and planning for robots, parsing multi-level spatial queries (object, room, floor), integrating Voronoi-based motion graphs, and achieving real-world navigation in multi-floor or dynamic environments [2403.17846], [2410.06239].
- **Scene Synthesis**: Hierarchical graph representations coupled with LLMs and hierarchy-aware graph nets enable scene generation pipelines that enforce both physical constraints and semantic fidelity, outperforming flat LLM methods in both physical feasibility and user-alignment [2502.10675].
- **Collaborative Mapping**: Multi-agent systems such as CURB-OSG and OpenGraph scale HOV-SG principles to large-scale outdoor or urban settings, with dynamic fusion, zero-shot semantic labeling, and hierarchical memory-efficient map structures [2503.08474], [2403.09412].
- **Task Reasoning and Symbolic Querying**: HOV-SG enables LLM-driven parsing of tasks and subgoal reasoning over structured semantic graphs, supporting robust agent interaction in realistic settings [2602.02456], [2412.19021].

## 7. Limitations, Open Challenges, and Future Research

Significant open issues remain:
- **Hierarchy/Cluster Design**: Manual super-category taxonomies or static entity clustering risk misrepresenting semantic nuance; adaptive, data-driven clustering remains an area for improvement [2303.06842], [2412.19021].
- **Relation Diversity and Grounding**: Region-aware prompt mining via LLMs introduces factual or diversity limits, especially in long-tail compositional scenarios [2412.19021].
- **Implicit Bias and Noise Propagation**: Latency and noise emerge from reliance on foundation models; hallucinations or category under-specification can propagate through the graph hierarchy [2403.09412].
- **Learning-Based Refinement**: Many current HOV-SG frameworks use frozen vision–language backbones; end-to-end fine-tuning, contrastive objectives on open-vocab relation/attribute heads, and GNN-based graph refinement are promising directions [2303.06842], [2602.02456], [2403.09412].
- **Real-Time Constraints and Planner Integration**: LLM-based planning logic is frequently offboard and incurs latency; on-device quantized LLM integration and closed-loop fusion are critical for scaling to high-frequency adaptive interaction [2410.06239].

Ongoing work targets graph-neural extensions, richer affordance and functional edges, continual learning across agents, and extension to 4D spatiotemporal HOV-SG structures.

---

In summary, Hierarchical Open-Vocabulary Scene Graphs constitute a unifying formalism with demonstrated superiority in open-world vision-language tasks, embodied reasoning, and scalable robot scene understanding. By integrating multi-level abstraction, open-vocabulary semantics, and recent advances in LLMs and VLMs, HOV-SG frameworks are foundational components for the next generation of interpretable, robust, and generalizable perception and reasoning systems [2403.17846], [2602.02456], [2412.19021], [2502.10675], [2409.10350], [2410.06239], [2303.06842], [2403.09412], [2503.08474], [2504.13153].

Source: https://www.emergentmind.com/topics/hierarchical-open-vocabulary-scene-graphs-hov-sg