Abstract reasoning from minimal examples in frontier foundation models
Determine learning principles and model designs that enable frontier foundation models (e.g., GPT-5 and Grok 4) to correctly infer structured transformation rules from a handful of input–output matrix pairs and generalize these rules to novel test matrices in the Abstraction and Reasoning Corpus (ARC-AGI).
References
Abstract reasoning from minimal examples remains a core unsolved problem for frontier foundation models such as GPT-5 and Grok 4. These models still fail to infer structured transformation rules from a handful of examples, which is a key hallmark of human intelligence.
— Think Visually, Reason Textually: Vision-Language Synergy in ARC
(2511.15703 - Zhang et al., 19 Nov 2025) in Abstract (page 1)
A likely explanation is that many test tasks are out-of-distribution, though this cannot be confirmed without explicit test-rule annotations.
— Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks
(2609.21181 - Deliège et al., 18 Sep 2026) in Section 4.2, subsection “Qualitative ‘out-of-distribution’ evaluation of implicit rule induction,” paragraph “Test performance predictability”