Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

DALL-E 2 Fails to Reliably Capture Common Syntactic Processes (2210.12889v2)

Published 23 Oct 2022 in cs.CL and cs.CV

Abstract: Machine intelligence is increasingly being linked to claims about sentience, language processing, and an ability to comprehend and transform natural language into a range of stimuli. We systematically analyze the ability of DALL-E 2 to capture 8 grammatical phenomena pertaining to compositionality that are widely discussed in linguistics and pervasive in human language: binding principles and coreference, passives, word order, coordination, comparatives, negation, ellipsis, and structural ambiguity. Whereas young children routinely master these phenomena, learning systematic mappings between syntax and semantics, DALL-E 2 is unable to reliably infer meanings that are consistent with the syntax. These results challenge recent claims concerning the capacity of such systems to understand of human language. We make available the full set of test materials as a benchmark for future testing.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (3)
  1. Evelina Leivada (7 papers)
  2. Elliot Murphy (11 papers)
  3. Gary Marcus (13 papers)
Citations (35)
X Twitter Logo Streamline Icon: https://streamlinehq.com