CreativityBench: Affordance-Based Tool Use
- CreativityBench is a large-scale benchmark evaluating affordance-based creative tool use through reasoning on objects, parts, and their physical attributes.
- It utilizes a structured knowledge base with detailed household scene annotations to generate tasks that demand precise part selection and mechanism grounding.
- The benchmark measures both objective tool selection and explanation quality, highlighting discrepancies between plausible entity choices and correct part-level reasoning.
Searching arXiv for the benchmark and closely related variants to ground the article in the current literature. CreativityBench is a large-scale benchmark designed to rigorously evaluate a distinct facet of intelligence in LLMs and agents: affordance-based creative tool use, defined as the ability to discover and justify non-obvious yet physically plausible solutions under constraints by reasoning at the level of objects, parts, attributes, and mechanisms (Qian et al., 6 Apr 2026). Rather than testing canonical tool use or abstract puzzle solving, it evaluates whether a model can infer and exploit latent capacities emerging from part-level attributes such as rigidity, sharpness, elasticity, and geometry. In adjacent literature, closely named systems include AGC-Bench, described as a consolidated, psychometrically calibrated benchmark for artificial general creativity (“CreativityBench”), and MM-CreativityBench, which extends affordance-grounded creative tool use to visually rich, interactive settings (Beaty et al., 1 Jul 2026, Qian et al., 25 May 2026).
1. Nomenclature and position in the literature
The name “CreativityBench” is not unique in recent arXiv literature. It most directly denotes the affordance-based benchmark introduced in “CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing” (Qian et al., 6 Apr 2026), but neighboring work uses similar labels for different constructs.
| Benchmark | Primary target | Distinguishing feature |
|---|---|---|
| CreativityBench (Qian et al., 6 Apr 2026) | Affordance-based creative tool use | Part-level grounding over objects, attributes, and mechanisms |
| AGC-Bench (Beaty et al., 1 Jul 2026) | Artificial general creativity | 78 datasets across six domains with Judge Response Theory |
| MM-CreativityBench (Qian et al., 25 May 2026) | Creative physical intelligence | Interactive scene→entity→part inspection in multimodal environments |
| CreativeBench (Wang et al., 12 Mar 2026) | Machine creativity in code generation | Unified metric defined as quality × novelty |
This naming overlap reflects a broader fragmentation in creativity evaluation. CreativityBench (Qian et al., 6 Apr 2026) is unusual in that it isolates a narrow but operationally important capability: creative tool repurposing grounded in physical affordances, rather than general ideation, humor, writing, or code generation. A plausible implication is that it occupies a middle position between open-ended creativity benchmarks and embodied planning benchmarks: more grounded than generic ideation, but more explicitly creative than standard affordance or commonsense datasets.
2. Conceptual basis: affordance-grounded creative tool use
CreativityBench isolates “affordance-based creative tool use,” the ability to infer and exploit an object’s affordances—the action possibilities enabled by its physical and state attributes—to achieve a goal in an unconventional yet physically plausible way (Qian et al., 6 Apr 2026). In canonical tool use, an object is used for its intended function; in affordance-based creative tool use, the model must identify latent capacities emerging from part-level attributes and repurpose the tool accordingly.
The benchmark draws inspiration from the notion of affordances in ecological psychology and adopts the robotics-style linkage between perception and action. Its central claim is that strong general reasoning does not automatically translate into mechanism-level grounding. Models often choose a plausible object, but fail to identify the correct part, the enabling attributes, and the physical mechanism needed to solve the task. This distinguishes CreativityBench from deduction-heavy reasoning tasks, which may reward planning or symbolic inference without requiring a grounded account of why a particular part can do a particular job (Qian et al., 6 Apr 2026).
This focus also differentiates CreativityBench from broader “general creativity” infrastructures. AGC-Bench aggregates creativity across brainstorming, problem solving, STEM, narrative, figurative language, and humor, and uses psychometric calibration to recover a single creativity factor over LLMs (Beaty et al., 1 Jul 2026). CreativityBench instead operationalizes one sharply defined slice of creativity: grounded repurposing under physical constraints.
3. Knowledge base and benchmark construction
At the core of CreativityBench is a structured affordance knowledge base built via a staged, LLM-assisted annotation pipeline (Qian et al., 6 Apr 2026). The KB spans eight everyday household scenes—Bathroom, Bedroom, Dining Room, Garage, Garden, Home Office, Kitchen, and Living Room—and contains 3,816 entities, 26,238 parts, 288,318 physical-attribute annotations, 124,972 state-attribute annotations, and 157,427 affordance annotations, for 570,717 total annotations.
The KB is formalized as a directed multigraph with , where denotes objects, parts, attributes, and affordances or uses (Qian et al., 6 Apr 2026). Affordances are represented together with use, environment, and recipient conditions: where is the core action, the use condition, 0 the environment condition, and 1 the recipient condition. The resource also records affordance typicality labels, ranging from Normal 0 to Emergency 1–5.
Task generation is reverse-engineered from the KB rather than written first and annotated later. A task 2 consists of a scenario 3, a candidate entity set 4, and a gold solution 5, where the model must recover the correct entity, the correct part, and the correct affordance under scenario constraints (Qian et al., 6 Apr 2026). In the reported run, this process produced approximately 14K grounded tasks, specifically 14,280. Gold affordances are sampled using semantic clustering; scenarios are synthesized as first-person narratives; candidate golds are rejected if another part or entity affords a strictly preferable solution; and distractors are chosen to control difficulty via distractor count and affordance similarity.
Scaling is achieved with GPT-5.2 under strict JSON schemas and consistency checks. Quality assessment on a 0.1% sample reports approximately 98% pass under automatic LLM-judge checks and 95% under human review, while a separate human study of gold-solution persuasiveness reports average inter-annotator agreement of 63.0% (Qian et al., 6 Apr 2026).
4. Evaluation protocol and metrics
Each evaluation instance presents a scenario together with full descriptions of entities present in the scene, including part-level attributes. The model must select the correct entity and part, filter distractors, and generate a how-to-use explanation grounded in physical attributes and conditions (Qian et al., 6 Apr 2026). Two objective tool-usage metrics are central:
6
7
8
9
Here, 0 is the Gold Correct Rate and 1 is the Entity Correct Rate. For gold-correct predictions, an LLM-as-judge scores explanation quality on a 1–5 scale for use-condition coverage, environment-condition coverage, recipient-condition coverage, physical grounding, action feasibility, and prediction correctness (Qian et al., 6 Apr 2026).
The benchmark’s main metrics are deliberately split between objective selection and judged explanation. This separation matters because a model may name the correct entity while missing the correct part, or it may propose an action that sounds plausible while remaining mechanically invalid. A plausible implication is that CreativityBench measures grounded creativity as a layered capability: plausible object choice, precise part selection, constraint satisfaction, and mechanism-level justification must all align before a response counts as successful.
5. Empirical findings and failure structure
Across ten evaluated LLMs, the average Gold Correct Rate is 0.1910 and the average Entity Correct Rate is 0.5149, representing a greater than 60% drop from entity-level plausibility to gold-level correctness (Qian et al., 6 Apr 2026). Average physical grounding is 3.2003 and average action feasibility is 3.5860, indicating that actions are often judged more plausible than well grounded. Recipient constraints are most frequently overlooked, with average 2 coverage of 2.8026, lower than 3 and 4.
Qwen3-32B achieves the highest Gold Correct Rate, 0.2588, and the highest Entity Correct Rate, 0.6246, outperforming GPT-5.2 on tool selection, even though GPT-5.2 has higher judged feasibility and correctness once it is on track (Qian et al., 6 Apr 2026). Scaling within model families shows saturation: gains are substantial at small sizes but diminish at larger scales. Inference-time interventions also help only marginally. Higher temperature often harms smaller models via hallucinated entities and parts; structured Chain-of-Thought yields small improvements in some open-source models and slight declines in some GPT-family models; and interactive mode consistently reduces performance because models inspect too few entities before committing.
Error analysis identifies a strong pattern. Gold solutions are overwhelmingly preferred, with a greater than 95% win rate in comparisons, indicating that wrong-tool predictions are rarely close substitutes (Qian et al., 6 Apr 2026). When the entity is correct but the part is wrong, judged action feasibility drops by approximately 0.6 absolute. Failures are dominated by physical invalidity, especially affordance mismatch and hallucinated affordances; practicality and safety constraints are secondary contributors. Typical mistakes include choosing the right object but the wrong part, ignoring state attributes, skipping recipient constraints, and proposing plausible-sounding yet mechanically invalid actions.
These findings support the paper’s central argument: present models often retrieve enough commonsense to identify an object class, but they do not reliably sustain the part-level grounding needed for creative affordance discovery.
6. Multimodal extension and related benchmark ecosystems
A direct multimodal extension appears in MM-CreativityBench, introduced as a benchmark for “creative physical intelligence” in visually rich, physically constrained environments (Qian et al., 25 May 2026). It retains the entity–part–affordance logic of CreativityBench, but makes evaluation interactive: the agent begins from an environment image, can request entity views, then part views, and must return structured actions of the form inspect_entity, inspect_part, or answer. The test set contains 333 tasks, 974 entities, and 6,344 parts; the training set contains 868 tasks, 1,498 entities, and 10,080 parts (Qian et al., 25 May 2026).
Base multimodal models remain weak on gold-correct performance. For example, GPT-5.4 reaches Gold Correct 0.192 and Entity Correct 0.435, while Qwen3-VL-32B reaches Gold Correct 0.240 and Entity Correct 0.447 (Qian et al., 25 May 2026). The gap between correct entity and correct part therefore persists in the visual setting. MM-CreativityBench then introduces affordance-grounded alignment using supervised fine-tuning and turn-level Direct Preference Optimization. On Qwen3-VL-4B, Gold Correct improves from 0.156 to 0.417; on Qwen3-VL-8B, from 0.192 to 0.393, while turns decrease substantially (Qian et al., 25 May 2026). Error category A, physical or functional invalidity, drops from 45.3% to 25.8% for Qwen3-8B after alignment.
The broader creativity-benchmark landscape provides complementary views. AGC-Bench studies artificial general creativity across 78 datasets and reports that a single factor 5 explains 81.5% of variance across six creativity domains (Beaty et al., 1 Jul 2026). “CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges” focuses on code generation and defines creativity as the product of quality and novelty (Wang et al., 12 Mar 2026). CREATE evaluates associative creativity through multi-hop concept paths in parametric knowledge (Wadhwa et al., 10 Mar 2026). CreativityBench is narrower than these efforts, but also more tightly grounded in physically plausible tool repurposing.
7. Limitations, implications, and resources
CreativityBench is text-only in its original form and household-focused in scope (Qian et al., 6 Apr 2026). Its single-gold solution design is explicitly a methodological choice for rigorous measurement, even though real-world creativity may admit multiple viable solutions. Automatic judging remains substantial, and while prompts are strict and quality-checked, some subjectivity persists. The benchmark also inherits potential annotation bias from its LLM-assisted construction pipeline.
These constraints do not weaken its main contribution so much as delimit it. CreativityBench does not attempt to measure creativity as an unrestricted, global trait. Instead, it treats creative tool use as a structured reasoning problem in which novelty must remain physically grounded. This suggests a useful decomposition of creative competence in agents: general reasoning, object plausibility, part selection, attribute grounding, and mechanism justification can fail independently.
Resources are released publicly, including the code repository and project page, with data, KB schemas, task generation scripts, evaluation prompts, and instructions for running evaluations (Qian et al., 6 Apr 2026). MM-CreativityBench likewise releases code and data for interactive multimodal evaluation (Qian et al., 25 May 2026). Together, these benchmarks establish affordance-grounded creative tool use as a reproducible subfield of creativity evaluation rather than an anecdotal capability.