FP-LGN: Feature Pyramid Likelihood Grounding
- The paper presents FP-LGN as a probabilistic spatial grounding model that estimates likelihoods for human and robot observations using Bayesian fusion.
- It employs a feature pyramid network to extract multi-scale map features, enabling context-sensitive grounding of spatial relations with uncertainty.
- A three-stage curriculum learning strategy is introduced to progressively learn from synthetic data to human-collected data, achieving competitive negative log-likelihood performance.
Searching arXiv for the specified paper to ground the article in the primary source. The Feature Pyramid Likelihood Grounding Network (FP-LGN) is a probabilistic spatial language grounding model introduced for Bayesian fusion of human and robot observations. It estimates a likelihood over map locations for a parsed spatial relation, rather than assigning a deterministic label or region, thereby representing the uncertainty inherent in human spatial semantics. In the formulation reported in "Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations" (Sitdhipol et al., 26 Jul 2025), FP-LGN learns how spatial relation likelihoods adapt to contextual map geometry from data, with the explicit objective of supporting recursive Bayesian estimation in human-robot collaborative sensing.
1. Problem formulation and probabilistic role
FP-LGN addresses spatial language likelihood grounding. Given a map , a reference landmark , a spatial relation extracted from language, and a candidate target location , the model estimates
This formulation grounds expressions such as “near the building,” “in front of Building 5,” “at the entrance,” and “far from the building” relative to actual map context and landmark geometry (Sitdhipol et al., 26 Jul 2025). The target of inference is not a categorical attachment between phrase and landmark alone, but a location-conditioned likelihood over the environment.
The probabilistic role of this quantity is central. In the paper’s human-robot collaborative sensing setting, a robot receives robot sensor measurements and human natural-language observations . Sensor fusion in a recursive Bayesian framework requires each observation source to be expressed as a likelihood. For sensors, this is represented by ; for language, an uncertainty-aware system requires a corresponding human observation likelihood (Sitdhipol et al., 26 Jul 2025). This requirement motivates FP-LGN’s design as a likelihood estimator rather than a deterministic grounding module.
The LLM is integrated through recursive Bayesian updates. After prediction, robot measurements are incorporated via
and human language observations are fused analogously through
0
For a target 1 with a set of parsed spatial observations 2, the full language likelihood is factorized as
3
Within this framework, FP-LGN is the component that estimates the per-relation factor (Sitdhipol et al., 26 Jul 2025).
2. Motivation and distinction from rule-based grounding
The motivating problem is that prior approaches to spatial language grounding for robotics often relied on expert-designed assumptions about the geometric semantics of relations. The paper identifies dependence on predefined likelihood equations, parametric preposition models, random-set models, and manually specified geometric relationships as key limitations (Sitdhipol et al., 26 Jul 2025). Such methods can incorporate landmark geometry, but their adaptation remains constrained by expert-defined forms rather than learned directly from human judgments.
A second limitation is the handling of aleatoric ambiguity. Spatial expressions such as “near,” “in front of,” and “by” are not interpreted identically across annotators. A deterministic predictor or a classifier optimized for point estimation does not directly model this semantic variability. FP-LGN is therefore positioned as a model that learns a full likelihood distribution reflecting uncertainty and ambiguity in human spatial semantics, rather than merely predicting a best region.
The paper’s central claim is that FP-LGN is the first model in this line of work to fully learn how spatial language likelihood should adapt to contextual map geometry, rather than depending on hand-designed rules about landmark shape, entrances, boundaries, or reference directions (Sitdhipol et al., 26 Jul 2025). This claim should be read specifically within the paper’s line of work on probabilistic grounding for Bayesian fusion. A plausible implication is that the method is intended as a replacement for manually engineered relation likelihood templates in robot perception pipelines, not merely as an incremental improvement in visual-language grounding.
The significance of this shift is methodological. If language is to function as an observation source in Bayesian estimation, it must supply a probabilistic observation model. FP-LGN’s contribution is therefore not only representational but inferential: it allows human language to participate in recursive posterior updates without collapsing semantic ambiguity into a hard symbolic fact.
3. Architecture and representation
FP-LGN consists of three main components: Map Feature Extractor and Map Encoder, Spatial Relational Encoder, and Text and Map Embedding Interaction (Sitdhipol et al., 26 Jul 2025). An upstream parsing module extracts tuples 4 from free-form human language, after which the grounding model operates on a single relation-landmark pair.
The effective inputs for grounding one spatial relation are the map image, the queried location 5, the reference landmark 6, and the spatial relation 7. The output is the estimated probability
8
When evaluated across many candidate cells or locations, this output becomes a dense likelihood map over the environment (Sitdhipol et al., 26 Jul 2025).
The distinctive architectural choice is the use of a Feature Pyramid Network (FPN) as the map feature extractor. The rationale given is that spatial language grounding requires geometric landmark features at multiple levels of detail: fine-grained structure for boundaries, inside/outside distinctions, edges, and entrances, and coarser context for broader relations such as “near” or “far from.” According to the paper, map image features are extracted through the FPN; ROI Pooling is applied on feature map layers; average pooling is then applied to the ROI outputs to reduce dimensionality; pooled features are passed through dense layers in the map encoder; and the resulting encoded map embedding is concatenated with the relation embedding (Sitdhipol et al., 26 Jul 2025). The figure caption further notes that the queried location is included through ROI pooling, so the model extracts features around the candidate location from the pyramid features.
On the language side, the spatial relation 9 is represented as a one-hot vector and encoded by fully connected layers into a learned relation embedding. The paper states that this simple representation is effective and scalable across variations of spatial relations, and that keeping the relation embedding independent of specific map locations allows multiple relations to overlap semantically at one place; for example, a location can be both “near” and “in front of” a landmark (Sitdhipol et al., 26 Jul 2025).
The interaction module concatenates the map embedding and relation embedding, forms a dense fused representation, and applies a final layer with sigmoid activation to predict the likelihood. The model is therefore a discriminative neural network trained to output a probability estimate for whether a relation holds at a given map location under the given map context. The paper does not report the exact number of FPN levels, channel dimensions, hidden layer widths, kernel sizes, ROI pooling resolution, embedding sizes, or parameter counts, and these architectural details are absent from the manuscript (Sitdhipol et al., 26 Jul 2025).
4. Grounding mechanism and uncertainty modeling
FP-LGN grounds each parsed relation-landmark pair by estimating
0
so grounding occurs at the level of a candidate target location relative to a named landmark in a specific map (Sitdhipol et al., 26 Jul 2025). The experiments and figures indicate that the model is used to generate dense likelihood distributions over map locations, represented as heatmaps over map images. The output should therefore be understood as a probability or likelihood over map cells or sampled spatial locations, not over a small discrete set of object candidates.
The grounding mechanism depends jointly on the geometric structure of the reference landmark, surrounding context such as nearby buildings, roads, and entrances, and the semantics of the spatial relation. The qualitative results reported in the paper show several behaviors. For “at,” “near,” and “far from,” the learned likelihood often follows landmark contours. For small concave buildings, the learned “near/far” behavior may follow a convex hull rather than the literal concavity, which the paper states better matches human labeling. For “in front of,” the model can produce multimodal likelihoods, particularly for buildings with multiple entrances, with high-probability regions in front of each entrance (Sitdhipol et al., 26 Jul 2025). This suggests that the model learns a context-conditioned semantics distribution rather than a fixed relation template.
The paper explicitly frames the uncertainty of linguistic interpretation as aleatoric uncertainty, arising from “inherent semantic ambiguity and variability in the interpretation of spatial expressions within the human population” (Sitdhipol et al., 26 Jul 2025). FP-LGN is therefore trained as a probability estimator that outputs a predicted likelihood 1 to approximate a true semantic distribution 2. The objective is stated as minimizing
3
which the paper notes corresponds to minimizing expected negative log-likelihood (NLL) over data sampled from 4. The manuscript does not provide a fully expanded sample-wise NLL equation with explicit dataset summation indices; the KL/NLL relationship is stated conceptually (Sitdhipol et al., 26 Jul 2025).
This distinction from ordinary classification is central to the method’s interpretation. A classifier would learn decision boundaries and optimize for class prediction, whereas FP-LGN is designed to capture repeated human judgments as a distribution. The paper’s training curriculum is organized around that distinction: first learning decision structure, then synthetic uncertainty, and finally human aleatoric variability.
5. Curriculum learning, data, and supervision
A three-stage curriculum learning strategy is a core training contribution. The paper argues that the task is harder than standard classification because the model must learn relevant geometric context from maps, uncertainty-aware semantics, and real human ambiguity from sparse labels. Without staged training, optimization is described as unstable or suboptimal (Sitdhipol et al., 26 Jul 2025).
In Stage 1, FP-LGN is trained like a regular classifier on synthetic data. The purpose is to learn basic spatial decision boundaries, useful map feature extraction for varied geometric structures, and point estimation rather than full uncertainty. This provides an initialization for recognizing map geometry and relation patterns. In Stage 2, the synthetic data generation process is modified so that synthesized labels include uncertainty; repeated sampling is used at each training input value, increasing data density needed to capture aleatoric ambiguity. The purpose is to move from hard classification to distribution learning and teach the model that the same map/location relation can admit uncertain labels. In Stage 3, the model is fine-tuned on real human-collected data to adapt from synthetic uncertainty to true human semantic variability and align the learned likelihood with crowd judgments (Sitdhipol et al., 26 Jul 2025).
The reported optimization details are specific: Adam optimizer, initial learning rate 5, StepLR scheduler with step size 10 and decay factor 0.6, and early stopping patience of 20 epochs (Sitdhipol et al., 26 Jul 2025). The paper does not provide the number of epochs per curriculum stage, synthetic dataset size, exact synthesis procedure, or stage transition criteria.
The human grounding dataset used for likelihood evaluation was collected through Prolific. It comprises 35 map regions sourced from OpenStreetMap (OSM) in Bangkok, Thailand and Washington, DC, USA, including elements such as buildings, entrances, and streets (Sitdhipol et al., 26 Jul 2025). Workers were shown one map at a time with a queried location, a reference landmark, and surrounding context, and judged whether spatial relations correctly described the location using Yes/No responses. The dataset used 56 crowdworkers. The paper states that no majority voting was used; instead, two rejection mechanisms were applied: worker-level exclusion for consistently poor labeling and individual-point rejection for clearly invalid labels. This preserves semantic variability while filtering obvious annotation errors.
The dataset covers ten common spatial relations from prior work, including “at,” “next to,” “in front of,” “by,” “near,” “far from,” “close to,” “beside,” “around,” and “alongside” (Sitdhipol et al., 26 Jul 2025). The reported split is 21 map regions, 2,404 locations for training and 14 map regions, 2,782 locations for testing. Training augmentation used random flips and random rotations. The conceptual ground truth is the underlying human semantic distribution 6, approximated through human labels, but the paper does not provide a detailed mathematical description of how multiple Yes/No labels are converted into a target probability for each location-relation pair.
6. Empirical evaluation and downstream search performance
The evaluation reported in the paper has two parts: likelihood grounding quality and downstream human-robot collaborative sensing (Sitdhipol et al., 26 Jul 2025). For likelihood grounding, the main metric is mean Negative Log-Likelihood (NLL) on unseen human-generated evaluation data, used as a proxy for information loss under the objective of minimizing 7.
The baselines are Expert, C-LGN, and Chance. The Expert baseline is a rule-based likelihood model from prior work with parameters optimized by maximum likelihood on the same training data. C-LGN is an ablation baseline that replaces the FPN with ResNet34. Chance is a random likelihood from a uniform distribution (Sitdhipol et al., 26 Jul 2025).
| Model | Mean NLL | SD |
|---|---|---|
| FP-LGN | 0.384 | 0.676 |
| Expert | 0.387 | 0.881 |
| C-LGN | 0.532 | 0.538 |
| Chance | 1.015 | 1.012 |
These results support two claims made in the paper. First, FP-LGN matched expert-designed rules in mean NLL, with no significant difference in mean NLL between FP-LGN and Expert (8). Second, FP-LGN demonstrated greater robustness, as indicated by lower NLL standard deviation than the Expert model: 0.676 versus 0.881 (Sitdhipol et al., 26 Jul 2025). FP-LGN also significantly outperformed C-LGN in mean NLL (9), and all tested models significantly outperformed Chance (0). The paper interprets the low SD of C-LGN not as superior reliability but as reflecting repeated failures on similar patterns.
The ablation result is important because it isolates the value of the feature pyramid. Replacing the FPN with ResNet34 degraded mean NLL from 0.384 to 0.532, indicating that multi-resolution map encoding is important for relations requiring fine-grained map resolution, including inside-building judgments, near-edge distinctions, and geometry-sensitive semantics (Sitdhipol et al., 26 Jul 2025).
The downstream application is a simulated target search task with 150 search scenarios on maps from three cities in Thailand. Robot and target positions were randomized; a human monitored via surveillance cameras and reported observations in natural language; the robot had a 360° camera; the target detector ran at 1 Hz with true positive and true negative rates of 0.8; detection range was 25 m; security cameras had 45° fixed cone FoV; robot speed was 1 m/s; and planning was performed toward the current MAP target estimate using A* (Sitdhipol et al., 26 Jul 2025). Human language had 83.0% subject-predicate structure and 17.0% existential structure; common relations included “in front of,” “near,” “close to,” “beside,” “next to,” “around,” and “alongside”; 32.5% of sentences contained at least one typo; and parser accuracy was 97.7%.
Four modes were evaluated: human-robot, robot-only, human-only, and uninformed. The success-rate results indicate that human-robot reached 100.0% success within 4043 steps, robot-only required 8300 steps to reach 100.0%, human-only plateaued at 94.0% after 4807 steps, and robot-only plateaued at 96.0% after 4872 steps (Sitdhipol et al., 26 Jul 2025).
| Mode | Mean search steps | SD |
|---|---|---|
| human-robot | 1054 | 1065 |
| robot-only | 2021 | 1610 |
| human-only | 1804 | 2273 |
| uninformed | 8411 | 18739 |
The collaborative mode reduced mean search steps by 47.8% relative to robot-only, 41.6% relative to human-only, and 87.4% relative to uninformed, with all reductions statistically significant at 1 (Sitdhipol et al., 26 Jul 2025). These results establish that FP-LGN’s likelihood output is not merely interpretable in isolation; it is sufficiently well-formed to support recursive uncertainty-aware fusion with physical sensor likelihoods in a search task.
7. Interpretation, limitations, and relation to prior work
FP-LGN is characterized in the paper by five main innovations: likelihood-based grounding rather than deterministic grounding, full adaptation to contextual map geometry, feature pyramid map encoding, explicit aleatoric uncertainty modeling, and three-stage curriculum learning (Sitdhipol et al., 26 Jul 2025). Its principal difference from previous work is that it learns the map-geometry-to-likelihood transformation itself, whereas prior approaches typically predefined relation likelihood formulas, fitted hand-designed parametric models, or used learned components only for subproblems such as frame-of-reference inference while keeping relation likelihood hand-specified.
Several limitations are also explicit or implicit in the manuscript. Real human data are sparse, which is one reason the curriculum is needed. The parser is not the main focus of the contribution, although the paper reports using LLaMA 2 7B successfully. Detailed uncertainty diagnostics such as calibration metrics are not reported; the probabilistic claims are instead supported through NLL, NLL variance, and downstream Bayesian fusion performance (Sitdhipol et al., 26 Jul 2025). Architecture details remain under-specified, including exact FPN configuration, embedding dimensions, layer widths, batch size, and epoch counts. Negative language examples such as “not in front of Building 8” appear in the collaborative task, but the exact mathematical treatment of negation is not given.
These omissions matter for interpretation. The absence of explicit calibration metrics means that FP-LGN’s probabilistic quality is demonstrated indirectly rather than through formal calibration analysis. The missing implementation details limit exact reproducibility. Likewise, the lack of a mathematical account of negation means that the handling of negative statements cannot be fully characterized from the published description alone. None of these points negate the reported empirical findings; rather, they delimit what can be concluded directly from the paper.
The broader significance of FP-LGN is that it recasts natural-language spatial observations as a principled observation source for Bayesian robotics. By estimating
2
and composing these factors into
3
the model allows human language to enter recursive Bayesian updates alongside sensor likelihoods (Sitdhipol et al., 26 Jul 2025). This suggests a general design principle for collaborative autonomy: spatial language is most useful not when forced into crisp symbolic facts, but when grounded as a context-sensitive likelihood that preserves human semantic ambiguity.