---
title: 'Object Permanence: Definition, Applications, and Research'
url: https://www.emergentmind.com/topics/object-permanence
type: topic
---

# Object Permanence: Definition, Applications, and Research

Object permanence is the capacity to represent an object as continuing to exist, retain its identity, and remain localizable when it is temporarily absent from direct sensory access. In visual computation, it concerns inference over hidden object state during occlusion, containment, carrying, out-of-view disappearance, or sensory interruption; in robotics and autonomous driving, it additionally concerns action-conditioned state maintenance, uncertainty management, and downstream planning. Contemporary formulations distinguish persistence from mere re-identification after reappearance: an online system must maintain a hypothesis before future evidence becomes available, whereas an offline system may use observations after the occlusion to reconstruct a globally consistent trajectory.

## 1. Psychological and developmental foundations

Piaget originally argued that object permanence develops relatively late and requires substantial sensorimotor interaction. Subsequent developmental studies reported that infants can reason about some occluded objects relatively early, while containment develops later. Computational work has consequently distinguished ordinary occlusion from containment: a hidden object may continue its own motion behind an occluder, whereas an object inside a moving container must be localized by following the container rather than extrapolating the target independently [2003.10469].

Controlled-rearing experiments with newborn chicks provide evidence that object permanence can be expressed without ordinary postnatal occlusion experience. In one experiment, chicks reared in a virtual environment in which objects never occluded one another nevertheless searched above chance for an object hidden behind a screen. Object-movement trials yielded $t(7)=6.52$, $p=.0003$, $d=2.31$, and invisible-displacement trials yielded $t(7)=4.10$, $p=.005$, $d=1.45$. In a second experiment, chicks exposed to thousands of hidden teleportation events continued to search at the location predicted by continuous motion rather than at the location favored by their rearing history. Chicks reared in the natural world yielded $t(3)=4.58$, $p=.02$, $d=2.29$, while chicks reared in the teleportation world yielded $t(3)=4.01$, $p=.03$, $d=2.00$ [2402.14641].

These findings are consistent with an early-developing continuity prior, possibly supported by prenatal spontaneous neural activity and early plasticity. They do not establish that a symbolic rule is genetically specified, nor do they prove that the relevant representation is formed prenatally. The experiments used small samples, virtual stimuli, and search behavior as an indirect measure. They also distinguish hidden violations, which did not reverse object permanence, from visibly discontinuous motion, which previous controlled-rearing work associated with impaired object permanence.

An alternative computational account treats object permanence as a product of general-purpose probabilistic program induction. In this formulation, a learner searches over compact transformations of discrete visual representations, including translation, union, intersection, and imputation. A program such as $\lambda\,(\mathrm{move}\ x\ n)$ can encode an object continuing through space and time, while latent states can be imputed during occlusion. The model learned stronger persistence when an object later reappeared than when it disappeared without reappearing, because the before-and-after sequence supplied stronger evidence for a continuing object [2309.07099]. This supports a computational-sufficiency argument: an innate object-specific module is not logically necessary to produce object-persistent behavior, although the program language and simplicity prior themselves provide substantial inductive bias.

## 2. Computational formulations and distinctions

Object permanence is not a single task. It includes several related but non-equivalent capabilities:

- **Visible-object localization**: estimating an object from currently available pixels.
- **Occluded-object prediction**: forecasting the position of an independently moving object with no visible target pixels.
- **Containment reasoning**: inferring that an object is inside or covered by another object.
- **Carried-object prediction**: following a hidden object through the changing state of a moving carrier.
- **Out-of-view persistence**: maintaining an object after camera motion or field-of-view exit.
- **Identity preservation**: reconnecting a reappearing object to its prior identity.
- **Amodal completion**: predicting the full extent of a partially or fully hidden object.
- **Uncertainty maintenance**: representing multiple plausible hidden states rather than a falsely precise single estimate.

The distinction between occlusion and containment is functional rather than purely visual. In ordinary occlusion, the system should continue reasoning about the target’s own motion. In containment, it should redirect spatial reference to a visible covering or containing object. In carrying, it must infer that the target’s location changes because the carrier changes location. The target may remain completely invisible throughout the transport [2003.10469].

Online object permanence differs from offline tracklet linking. An online system at time $t$ may use only observations through $t$ and must output a hidden-object hypothesis before reappearance. Offline tracking may use observations both before and after disappearance, together with map information, to associate fragmented tracklets and complete the missing interval [2012.08419; 2310.10372]. Re-identification restores identity retrospectively, but does not by itself provide the object’s location during the invisible interval.

A related distinction concerns existence and observability. A physically present actor may have no current sensor support. BeyondSight formalizes this distinction with an observability variable $o_t^i\in\{0,1\}$ and trains separate persistence and observability behavior. Its nuScenes-Permanence extension retains actors with zero LiDAR or radar support, interpolates temporary gaps, and extrapolates terminal states [2607.09138]. This formulation is particularly relevant to autonomous driving, where filtering out fully unobservable actors from labels can implicitly teach a model that “not observed now” means “does not exist.”

## 3. Learned representations and model architectures

### Localization and recurrent reference switching

OPNet formulates object permanence as frame-level bounding-box prediction. Its first recurrent module, the “Who to track?” module, assigns attention over detected object slots. Its second recurrent module, the “Where is it?” module, predicts the target box from the selected object representation and temporal history. The first module can attend to the target when visible, a static covering object during containment, or a moving carrier during carrying. The model is trained primarily through an $L_1$ localization loss rather than explicit supervision of containment relations [2003.10469].

On LA-CATER, the full OPNet obtained mean IoU values of 88.89 for visible frames, 78.83 for occluded frames, 76.79 for contained frames, 56.04 for carried frames, and 81.94 overall under detector-based perception. With perfect perception, the carried score increased to 76.42, demonstrating the importance of object identity and localization quality. In the original CATER snitch task, OPNet achieved 74.8% accuracy and an $L_1$ distance of 0.54 using a $6\times6$ location grid.

### Recurrent spatial memory and forecasting

PermaTrack extends CenterTrack from frame pairs to arbitrary-length online sequences with a convolutional gated recurrent unit. Its spatial memory maintains a distributed representation of previously observed objects, including objects currently invisible. Separate heads predict object centers, box sizes, displacement, and visibility. Invisible predictions are retained internally for identity maintenance rather than emitted as ordinary visible detections [2103.14258].

PermaTrack uses the Parallel Domain synthetic dataset, which provides amodal object annotations, persistent identities, visibility fractions, 3-D world coordinates, and camera parameters. Filtered supervision begins after an object has been visible for two consecutive frames. Three-dimensional constant-velocity pseudo-ground truth improved Track mAP to 67.0 compared with 65.7 for two-dimensional propagation and 61.2 for a CenterTrack heuristic. On KITTI, PermaTrack obtained car HOTA of 78.0 and person HOTA of 48.6; on MOT17 validation it obtained IDF1 of 67.0 with public detections and 68.2 with private detections.

Detecting Invisible People formulates online invisible-person detection as short-term trajectory forecasting under missing observations. The method combines constant-velocity dynamics, ego-motion compensation, monocular depth, and freespace reasoning. A predicted person is suppressed if it lies in visible freespace where it should have been detected, but retained if it lies behind the estimated nearest visible surface. On MOT-17, the full system improved occluded Top-5 F1 from 28.4 for DeepSORT to 39.8, an 11.4-point gain. The system’s test-set occluded Top-5 F1 was 43.4 on MOT-17 and 46.9 on MOT-20 [2012.08419].

### Object-centric memory and identity alignment

Objects-Align-Transition (OAT) uses MONet to decompose frames into object slots, then aligns those slots to a persistent memory using Memory AlignNet and Hungarian assignment. The persistent memory can contain more slots than are simultaneously visible; in the Playroom experiment, $K=10$ current slots were aligned to $M=12$ memory slots with $F=32$ latent features. A hidden object remains in its memory row while its current MONet slot is absent, and a later observation can be matched back to the same row [2103.04693].

OAT’s transition model predicts object-level latent changes rather than reconstructing pixels. This avoids background-dominance and ghosting, in which hidden objects fade into the background under pixel-level losses. In the Playroom environment, OAT obtained encoding ARI 0.62, unroll pixel error 0.0121, and unroll ARI 0.42, compared with unroll ARI 0.12 for the no-alignment model and 0.33 for OP3. In the CABI robotics dataset, OAT predicted the reappearance of a red object after approximately 51 time steps of full occlusion.

### Self-supervised temporal coherence and latent imagination

RAM, or Random walk Along Memory, avoids direct supervision of invisible positions. A ConvGRU produces sequence-conditioned spatial memory features, and a Markov walk propagates probability through a space-time graph of memory locations. Visible positions supervise the walk only when the object is visible; during occlusion, probability mass can flow through multiple plausible paths. The method therefore learns hidden trajectories from temporal coherence without constant-velocity assumptions or invisible-object labels [2204.01784].

On static-camera LA-CATER, RAM achieved mIoU of 91.7 for visible, 79.3 for occluded, 82.2 for contained, and 63.3 for carried frames. On LA-CATER-Moving, it achieved 90.0, 62.5, 55.1, and 51.8, respectively. It exceeded fully supervised OPNet on the moving-camera contained and carried categories by 15.1 and 21.0 mIoU points.

Loci-Looped combines slot-based object representations, autoregressive transition dynamics, pixel-space observations, and latent imagination. Each slot separates a Gestalt code for identity-related appearance from a positional code for location, size, and priority. Learned percept gates determine whether to fuse current observations or rely on latent predictions. During complete occlusion, an inner loop repeatedly applies the transition model without forcing the encoder to infer the object from absent evidence [2310.10372].

On ADEPT vanish videos, Loci-Looped achieved mean tracking error $2.6\pm2.7$, 96.6% successful tracks, and MOTA 0.84. Loci-Unlooped achieved 12.4 error, 7.4% successful tracks, and MOTA 0.76. In a violation-of-expectation experiment, a permanently vanished object generated higher slot-specific prediction error at the expected reappearance time, with $t(75)=1.69$, $p=.047$, and again when the occluder fell, with $t(75)=3.68$, $p<.001$. These effects are consistent with an internal prediction of continued existence, although they do not establish a human-like conceptual representation.

### Relation-aware segmentation and reconstruction

TCOW evaluates persistent, amodal, relation-aware tracking. Given a first-frame target mask, a causal model predicts target, frontmost occluder, and outermost container masks. The target mask is defined for every frame, including complete invisibility. Containment is determined from 3-D bounding-box overlap, with a containment threshold of 0.75; full invisibility is defined by an occlusion fraction of at least 0.95 [2305.03052].

On Kubric Random, full TCOW obtained target IoU 53.0 over all frames, 16.6 on fully invisible frames, occluder IoU 70.5, and container IoU 71.6. AOT, a strong video-object-segmentation model, obtained 41.3 target IoU and 6.8 invisible-target IoU after retraining. TCOW’s relation prediction was substantially stronger than its exact hidden-target reconstruction, indicating that identifying the surrounding occluder or container can be easier than predicting the target’s precise invisible mask.

PersistGS addresses object permanence in dynamic 3-D Gaussian Splatting. It decomposes a scene into static background Gaussians, per-object Gaussians, and collision meshes. During complete occlusion, the same object-level Gaussian representation is retained and transported along a differentiably simulated rigid-body trajectory. Friction and initial velocity are estimated from visible motion, while bounces, contact impulses, deceleration, and direction changes are supplied by rigid-body dynamics [2606.03479].

Across synthetic ball-fall, ball-bounce, and ball-roll scenes, PersistGS achieved mean PSNR 17.15 compared with 14.69 for constant velocity and 12.01 for no physics. Its mean trajectory RMSE was 0.600, compared with 1.515 for linear interpolation and 5.360 for constant velocity. A centroid silhouette loss reduced mean trajectory RMSE from 0.984 to 0.586, approximately a 40% reduction.

## 4. Action, relations, and physical state

Visual evidence alone is insufficient when an object undergoes invisible displacement. The same image sequence may be consistent with an object remaining stationary, entering a container, being grasped, or being transported. Action history supplies causal constraints that are unavailable from pixels alone [2110.00238].

Action-Aware Perceptual Anchoring (AAPA) maintains persistent symbolic anchors associated with physical object hypotheses. It uses detector outputs, temporal alignment, confidence, camera-motion compensation, occlusion rules, and attachment relations. An attachment relation $\mathrm{attached}(p_i,p_j)$ denotes a child object physically constrained to move with a parent. Such relations can arise from `pick-up`, `insert`, `screw-in`, or containment actions and can form a one-parent hierarchy. If a parent is itself attached to a higher-level parent, the system recursively propagates the child state through the hierarchy [2107.03038].

AAPA assumes that objects do not move autonomously, that state changes are caused by agent actions, that objects change gradually rather than teleporting, and that detector evidence should be temporally consistent. These assumptions make the system interpretable but limit it in environments involving autonomous motion or sudden unobserved interactions.

On LA-CATER with object detections, AAPA obtained overall mean IoU 84.66 and carried mean IoU 68.25, compared with 82.35 and 56.04 for OPNet. Its carried center-distance error was 4.65 pixels compared with 13.97 for OPNet. With perfect perception, AAPA obtained overall mean IoU 96.31 and carried mean IoU 82.54. In a gearbox assembly demonstration using a Universal Robot, attachment reasoning allowed hidden components to remain anchored to cases, subassemblies, or the robot hand while the camera viewpoint changed.

The action-aware framework also improved neural prediction. OPNet+AA increased carried mIoU from 77.08 to 87.81 under perfect perception, and reduced carried center error from 5.75 to 1.42 pixels. Under detector-based input, however, OPNet+AA sometimes failed because the action-identified parent was absent or misclassified in the detector output.

Object Permanence Filter (OPF) incorporates related principles into a particle filter for six-degree-of-freedom interactive-robot tracking. When direct measurements are absent, OPF uses a dynamics module, an occluder module, and an uncertainty module. It may extrapolate recent motion, transfer a virtual measurement from a likely occluder, branch across multiple occluder hypotheses, and enlarge covariance as the absence persists [2403.08231].

In simulation, OPF reduced general object-permanence translation error from 0.06289 for a standard particle filter to $0.01138\pm7.253\times10^{-4}$, and rotation error from 0.3870 to $0.06869\pm2.947\times10^{-3}$. In a sugar-dropping task, translation error was 0.02129 for OPF versus 0.05018 for the standard particle filter. A hardware sugar-dropping experiment succeeded in $8/10$ trials. OPF does not claim exact hidden-state recovery: covariance peaks during occlusion, and the robot can enter a safety mode when uncertainty exceeds a threshold.

## 5. Applications in robotics, autonomous driving, and 4-D perception

### Manipulation and error recovery

Audio-Visual Representations addresses a robotic drop-recovery problem. A Kinova Gen3 manipulator with an RGB-D wrist camera observes only part of a dropped cube’s trajectory, while a seven-channel microphone array records impact and bounce sounds. An audio encoder and trajectory encoder are fused before decoding a complete 135-point trajectory from a 65-point observed trajectory [2010.09948].

The method uses audio to infer impact timing, direction, bounce structure, and object-surface interaction, while partial vision corrects spatial offsets. It outperformed five baselines in the reported retrieval experiments and placed the robot sufficiently close for the cube to re-enter the wrist camera’s visual field, after which simple visual alignment supported grasping. The experiment used 1,404 valid trials, but its environment was narrow: one cube, one principal release height, one table configuration, and controlled drops.

### Autonomous driving

BeyondSight maintains actor queries across periods without current sensor evidence. A Temporal Prior Decoder propagates actor hypotheses without current image features; an Observation Decoder incorporates current evidence; and a Posterior Fusion Decoder reconciles both. Persistent queries are passed to motion prediction and planning rather than being used only for detection [2607.09138].

On standard nuScenes validation, BeyondSight improved SparseDrive mAP from 0.415 to 0.427, NDS from 0.526 to 0.536, and AMOTA from 0.372 to 0.401. Planning $L2_{\mathrm{avg}}$ decreased from 0.61 to 0.54. On nuScenes-Permanence, unobservable-actor mAP increased from 0.000 for SparseDrive to 0.249 for BeyondSight, while unobservable minADE improved from 0.615 to 0.479 and unobservable EPA from 0.000 to 0.285.

Offline Tracking with Object Permanence targets automated labeling of autonomous-driving datasets. It combines an online tracker, a future-aware Re-ID module, and a track-completion module conditioned on vectorized lane maps. The system associates pre-occlusion and post-occlusion tracklets and fills the missing interval with map-conditioned nonlinear motion. The supplied material identifies nuScenes as the evaluation setting but does not provide numerical metric tables or exact losses, association criteria, or ablations [2310.10372].

### Detection and tracking infrastructure

Integrated Object Permanence (IOP) inserts previous detections into the proposal stage of a two-stage detector. IOP lite concatenates predictions from the previous frame with current Region Proposal Network proposals; IOP with history uses multiple previous frames; and IOP with particles preserves multiple temporal hypotheses. The detector itself need not be retrained [2211.15505].

With Faster R-CNN, the baseline was 53.8 mAP. IOP lite improved average performance by 7.3 mAP points with approximately 3 ms overhead. IOP with 200 particles improved by 10.3 mAP points with approximately 79 ms overhead. The method’s main limitation is that it requires at least occasional initial detections and does not itself provide a full identity-preserving tracker.

### 4-D foundation models and visual memory

PersistBench evaluates whether 4-D reconstruction and video-generation models preserve an observed object after it leaves the input camera’s field of view. It uses approximately 24,000 raw $360^\circ$ clips processed into 2,000 high-quality paired sequences spanning ten object categories. An omniscient reference trajectory remains directed toward the target while the model receives an input trajectory in which the target becomes invisible [2609.20819].

The benchmark separates object permanence, motion continuity, and appearance preservation. Its permanence score requires both SAM2 trackability and VLM recognition of the same target. For dynamic objects, GEN3C achieved the highest reported invisible-segment permanence score, 87.06, while 4DGT achieved 3.89 and CUT3R 0.99. For static objects, Argus achieved 84.06 and GEN3C 83.97. Across models, visible-to-invisible degradation was systematic, supporting the conclusion that current systems often maintain short-term consistency without robust object-centric visual memory.

### Latent physical organization

Event-Conditioned Diagnostics studies whether passive object-state world models develop latent structure associated with kinematics, contact, and object permanence. GRU, Transformer-lite, and RSSM-lite models received object-level states and predicted future states over a fixed eight-frame horizon. Event-regime probes reached macro-F1 1.000 across reported architecture-seed cases. Field readouts showed kinematic dominance for free motion, increased contact emphasis for collision, and increased object-permanence emphasis for occlusion [2606.28455].

Projection-based causal field effects increased event-window prediction loss when field-aligned latent directions were suppressed. Collision-contact CFE averaged 0.00319, while hard-occlusion CFE averaged 0.00619. The occlusion effect was positive in all six cases but was not more damaging than random-subspace suppression, so it supports functional sensitivity to object-permanence-aligned structure without isolating a unique causal mechanism.

## 6. Evaluation, limitations, and unresolved questions

Evaluation methods reflect competing interpretations of object permanence. Bounding-box IoU and mIoU are appropriate for geometric localization, but they can obscure identity, uncertainty, and relation errors. Invisible-person detection uses Top-5 F1, Top-1 F1, IDF1, and occlusion-specific MOTA because hidden localization is intrinsically uncertain. TCOW evaluates target, occluder, and container masks separately. PersistBench separates permanence, continuity, and appearance, while nuScenes-Permanence separates observable and unobservable actors.

The central methodological difficulty is that exact hidden states may be ambiguous or unavailable. Directly supervising one invisible trajectory can impose an arbitrary target. RAM addresses this through probabilistic temporal walks. PermaTrack uses deterministic 3-D constant-velocity pseudo-ground truth. PersistGS uses differentiable rigid-body dynamics. BeyondSight uses completed trajectories with uncertainty-aware matching tolerances. These approaches are not interchangeable: pseudo-ground truth supplies a training target, probabilistic walks preserve alternatives, and physical simulation imposes a dynamical prior.

Common failure modes recur across systems:

- **Perception failure**: detector identity confusion, missing parent objects, imperfect segmentation, or incorrect depth ordering.
- **Identity swaps**: visually similar objects, crossing trajectories, collisions, and overlapping projected centers.
- **Stale persistence**: retaining an actor after it has exited the scene or permanently disappeared.
- **Motion drift**: constant-velocity or learned extrapolation becomes inaccurate during long gaps, abrupt turns, stops, or interaction.
- **Containment ambiguity**: nested containers, identical containers, transparent containers, and multiple plausible carriers.
- **Representation collapse**: pixel-level models allow hidden objects to fade, blur, or merge into the background.
- **Uncertainty miscalibration**: a single precise prediction is inappropriate when multiple hidden trajectories remain plausible.
- **Domain shift**: synthetic geometric scenes do not capture clutter, articulated bodies, transparent materials, deformable objects, camera motion, or open-world interactions.

The limitations of current evidence also constrain claims about “genuine” object permanence. Successful reappearance tracking can result from short-term motion correlation or post hoc association. High mask IoU does not establish identity preservation, and a VLM-recognized object may still follow an incorrect trajectory. Latent probes demonstrate decodability, but not necessarily causal use; projection interventions can disrupt correlated computations without identifying an isolated persistence mechanism. Behavioral experiments demonstrate robust expectations but do not uniquely determine their neural or computational implementation.

Several research directions follow from the combined results. Object-centric world models require explicit separation of existence, observability, identity, and state uncertainty. Containment and carrying require relational or action-conditioned representations rather than generic temporal memory. Robust systems should combine learned dynamics with physical priors, symbolic attachment relations, map constraints, or multimodal evidence such as audio. Benchmarks should include hidden-state references, nested containment, moving cameras, long and repeated occlusions, reappearance identity, uncertainty calibration, and downstream planning consequences. The distinction between visible tracking and hidden-state maintenance should remain explicit: object permanence is not simply better detection, longer temporal context, or successful re-identification, but the continued maintenance of a physically and relationally coherent object hypothesis when direct evidence is absent.

Source: https://www.emergentmind.com/topics/object-permanence