Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties (2306.15668v2)

Published 27 Jun 2023 in cs.CV, cs.AI, cs.GR, and cs.RO

Abstract: General physical scene understanding requires more than simply localizing and recognizing objects -- it requires knowledge that objects can have different latent properties (e.g., mass or elasticity), and that those properties affect the outcome of physical events. While there has been great progress in physical and video prediction models in recent years, benchmarks to test their performance typically do not require an understanding that objects have individual physical properties, or at best test only those properties that are directly observable (e.g., size or color). This work proposes a novel dataset and benchmark, termed Physion++, that rigorously evaluates visual physical prediction in artificial systems under circumstances where those predictions rely on accurate estimates of the latent physical properties of objects in the scene. Specifically, we test scenarios where accurate prediction relies on estimates of properties such as mass, friction, elasticity, and deformability, and where the values of those properties can only be inferred by observing how objects move and interact with other objects or fluids. We evaluate the performance of a number of state-of-the-art prediction models that span a variety of levels of learning vs. built-in knowledge, and compare that performance to a set of human predictions. We find that models that have been trained using standard regimes and datasets do not spontaneously learn to make inferences about latent properties, but also that models that encode objectness and physical states tend to make better predictions. However, there is still a huge gap between all models and human performance, and all models' predictions correlate poorly with those made by humans, suggesting that no state-of-the-art model is learning to make physical predictions in a human-like way. Project page: https://dingmyu.github.io/physion_v2/

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Hsiao-Yu Tung (9 papers)
  2. Mingyu Ding (82 papers)
  3. Zhenfang Chen (36 papers)
  4. Daniel Bear (3 papers)
  5. Chuang Gan (195 papers)
  6. Joshua B. Tenenbaum (257 papers)
  7. Daniel LK Yamins (4 papers)
  8. Kevin A. Smith (7 papers)
  9. Judith E Fan (4 papers)
Citations (8)
Github Logo Streamline Icon: https://streamlinehq.com

GitHub