---
title: Creativity Benchmark Evaluation
url: https://www.emergentmind.com/topics/creativity-benchmark
type: topic
---

# Creativity Benchmark Evaluation

Taken together, recent work uses the term *creativity benchmark* for evaluation frameworks that test whether AI systems can produce novel and appropriate solutions, generate multiple distinct outputs, or make non-obvious connections under domain-specific constraints. The literature no longer treats creativity as a single capability: benchmarks target combinational p-creativity, functional diversity, mathematical construction, real-world creative problem-solving, literary transfer, multimodal sensemaking, design evaluation, and marketing ideation, often with distinct scoring protocols and human-alignment requirements [1805.03720][2504.05228][2505.08744].

## 1. Historical emergence and scope

A major early reference point is the Creative Invention Benchmark (CrIB), a **2000-problem** benchmark distributed evenly across **five diverse domains**—Painting, Alien Language, Photobashing, Narrative, and Dessert Recipes—designed to evaluate *combinational p-creativity*, namely the ability to recombine existing knowledge into artifacts novel to the individual agent [1805.03720]. CrIB was presented as, to the authors’ knowledge, the first benchmark to test cross-domain combinational creativity, and it established several design principles that persist in later work: a limited initial knowledge base, domain-specific operators, hidden goals, and continuous feedback rather than direct target exposure [1805.03720].

Later benchmarks broadened both modality and construct. NoveltyBench evaluates whether language models can produce multiple *functionally distinct* responses rather than a single high-quality answer [2504.05228]. DeepMath-Creative shifts the focus to constructive mathematical ability through proofs and counterexamples [2505.08744]. CresOWLve evaluates creative problem-solving grounded in real-world knowledge rather than contrived riddles [2604.03374]. GuessBench moves into multimodal “creativity in the wild” by placing vision-language models in the role of guessers in Minecraft “Guess the Build” gameplay [2506.00814]. Creation-MMBench and CreBench target multimodal creative intelligence and human-aligned creativity assessment for MLLMs, while the Human Creativity Benchmark (HCB) evaluates model outputs in professional creative workflows rather than isolated prompts [2503.14478][2511.13626][2606.30561].

The field has also produced meta-benchmarks. AGC-Bench was built from a PRISMA-compliant systematic review that screened **3,101 papers**, identified **497 unique creativity benchmarks**, and onboarded **78 datasets** into a HELM-standardized harness [2607.01152]. This move from single-task evaluation to benchmark aggregation reflects a central conclusion of the literature: creativity is heterogeneous enough that isolated tests are insufficient for comparative model assessment.

## 2. Competing definitions of creativity

The benchmarks differ first in what they mean by creativity. In CrIB, the operative construct is *combinational p-creativity*: novelty arises by combining familiar components, and the novelty need only be new to the agent rather than to humanity at large [1805.03720]. In NoveltyBench, the emphasis is not artifact invention but *functional diversity*: two outputs are distinct only if seeing both provides genuinely new utility to a user [2504.05228]. In DeepMath-Creative, creativity is tied to *constructive* mathematical behavior—generating a proof or an explicit counterexample—rather than simply reaching a correct answer [2505.08744].

Other frameworks explicitly decompose creativity into dimensions. CreativityPrism organizes evaluation around **quality**, **novelty**, and **diversity**, spanning **9 tasks**, **3 domains**, and **20 task-specific evaluation metrics** [2510.20091]. C²-Eval distinguishes **convergent creativity**, where tasks admit constrained solutions, from **divergent creativity**, where tasks are open-ended, and evaluates both through **Usefulness, Originality, and Surprise (U-O-S)** [2510.04009]. NoveltyBench’s distinction between distinct$_k$ and utility$_k$ likewise separates mere multiplicity from user-relevant diversity [2504.05228].

Some work questions whether creativity should be reduced to a single latent trait at all. AGC-Bench reports that factor analysis across **83 models** and **6 creative domains** recovers a single creativity factor \(c\) explaining **81.5\% of variance**, while also showing that models rank differently across domains such as writing and scientific ideation [2607.01152]. HCB, by contrast, argues that evaluator disagreement in creative domains is not measurement noise but a meaningful signal; it therefore separates **convergence**, where professionals align around shared best practices, from **divergence**, where individual taste legitimately varies [2606.30561]. This is not merely a psychometric detail: it changes what a creativity benchmark is supposed to preserve.

## 3. Task families and benchmark construction

The construction of creativity benchmarks follows several recurring patterns: hidden-target reconstruction, constrained open-ended generation, real-world puzzle solving, workflow-based professional judgment, and multimodal sensemaking.

CrIB exemplifies hidden-target reconstruction. In each problem, the agent receives a restricted knowledge base and a domain-specific function such as `Paint`, `AddWord`, `Stamp`, or `Submit`, plus a `Clear` function and a scoring function returning feedback in \([0,1]\) [1805.03720]. The target artifact is unseen, so the system must infer it indirectly. CS\(^4\) uses a different anti-memorization strategy: it increases prompt specificity by controlling the number of synthesized story-writing constraints, ranging from **7** to **39**, so that models are less able to retell narratives from training data [2410.04197].

Constructive benchmarks in technical domains often use expert-authored items and manual grading. DeepMath-Creative contains **~179 problems**, about **60\% undergraduate-level** and **40\% master’s-level**, all **original, manually crafted by math experts** [2505.08744]. Each item uses a bidirectional format: if a proposition holds, provide a proof; if not, construct a counterexample [2505.08744]. CreativeBench, in code generation, distinguishes **CreativeBench-Combo** for combinatorial creativity from **CreativeBench-Explore** for exploratory creativity, using automated reverse engineering and self-play with constraint stacking [2603.11863].

Real-world creative problem-solving benchmarks deliberately move beyond synthetic prompts. CresOWLve is built from **2,061 questions** curated from “What? Where? When?” and annotates each puzzle with difficulty, knowledge domains, creative reasoning type, and cultural origin [2604.03374]. CREATE evaluates *associative creativity* through **931 curated queries** that require models to generate multiple factual paths connecting entities via their parametric knowledge [2603.09970]. LiveIdeaBench evaluates scientific idea generation from **1,180 keywords** spanning scientific domains, using minimal context rather than rich retrieval or paper-specific inputs [2412.17596].

Multimodal creativity benchmarks extend the same logic to image-grounded or workflow-grounded settings. GuessBench uses **1,500 images** and **2,000+ problems** from Minecraft “Guess the Build,” with static and dynamic settings and partially revealing hints [2506.00814]. Creation-MMBench contains **765 test cases** spanning **51 fine-grained creative tasks** with instance-specific evaluation criteria and role-based prompts [2503.14478]. CreBench operationalizes creativity from **idea** to **process** to **product**, with **12 indicators** scored on a five-point rubric [2511.13626]. HCB evaluates outputs across **five creative domains** and **three workflow phases**—ideation, mockup, and refinement—using professional judgments rather than crowd preferences [2606.30561].

| Benchmark | Task setting | Central construct |
|---|---|---|
| CrIB | 5 hidden-goal invention domains | Combinational p-creativity |
| NoveltyBench | Multiple LM outputs per prompt | Functional diversity |
| DeepMath-Creative | Proof / counterexample construction | Mathematical creativity |
| GuessBench | Minecraft build guessing | Multimodal sensemaking creativity |
| AGC-Bench | 78 standardized datasets | Artificial general creativity |
| HCB | Professional workflow evaluation | Convergence and divergence |

## 4. Scoring, calibration, and human judgment

Creativity benchmarks use markedly different scoring schemes because they are trying to isolate different failure modes. CrIB defines a normalized creativity score that subtracts the *Uncreative Max* baseline:
\[
Score = \frac{NScore_a - NScore_u}{400 - NScore_u}
\]
so that gains achievable without invention are discounted [1805.03720]. The purpose is explicit: improvement above the uncreative ceiling should indicate creative recombination rather than brute-force exploitation of the initial knowledge base [1805.03720].

NoveltyBench formalizes diversity at the set level. Its distinct$_k$ metric counts the number of functionally unique equivalence classes in \(k\) sampled generations, while utility$_k$ combines novelty and quality under a user-patience parameter \(p\), with only non-redundant generations contributing value [2504.05228]. CREATE follows a related set-based logic: its *creative utility* rewards larger sets of high-quality and diverse associative paths, with path quality determined by specificity and factuality and inter-path diversity determined from embedding distances [2603.09970]. CreativeBench makes the tradeoff explicit in another way by defining creativity as the product of **quality** and **novelty**, using executable code to separate creativity from hallucination [2603.11863].

Because many creative outputs are not fully autogradable, several benchmarks rely on LLM-as-a-judge or hybrid expert pipelines. Creation-MMBench combines unitary scoring for *Visual Factuality* with pairwise comparison against a baseline, and reports *Reward* and *Win Rate* under dual evaluation to reduce position bias [2503.14478]. The literary translation paired-task framework uses expert human annotations and UCP-based automatic scoring, with a creativity score defined by successful *Creative Shifts* minus unacceptable renderings over the total number of Units of Creative Potential [2604.18169]. CreBench evaluates alignment to human expert scores by Pearson correlation across idea, process, and product dimensions [2511.13626].

Judge bias is now a central methodological issue. AGC-Bench applies **Judge Response Theory (JRT)** to correct for judge severity and leniency and fine-tunes **AGC-Judge** on **48,299 JRT-corrected ratings** [2607.01152]. By contrast, the marketing-focused Creativity Benchmark reports that three LLM-as-judge setups show weak, inconsistent correlations with human rankings and judge-specific biases, leading the authors to conclude that automated judges cannot substitute for human evaluation [2509.09702]. HCB makes the related but stronger claim that collapsing professional disagreement into a single quality score discards actionable information about where models must be correct and where they should remain steerable [2606.30561].

## 5. Empirical regularities

Despite heterogeneity in task design, several empirical patterns recur. First, current models often perform substantially worse on creative requirements than on adjacent reasoning or retrieval tasks. NoveltyBench finds that even state-of-the-art language models generate *fewer than four* functionally distinct responses from **10 samples** for most diverse prompts, whereas humans regularly achieve **7–8** [2504.05228]. DeepMath-Creative reports that, even under lenient scoring, the best-performing model achieves merely **70\% accuracy**, primarily on basic undergraduate-level constructive tasks, with sharp decline on harder items and failure to provide substantive strategies for open problems [2505.08744]. CresOWLve shows a consistent creativity gap: models perform much better on factual questions than on creative ones, with drops on English from **-6.08\%** to **-17.21\%** depending on model [2604.03374].

Second, strong general ability does not reliably transfer to creativity. The literary translation paired-task framework finds that high source-text comprehension does not yield human-level translational creativity; only one model, Mistral-Large, comes close to the human En–Zh creativity score (**0.167** vs. **0.246**), and only three model–prompt combinations exceed **0.1** [2604.18169]. LiveIdeaBench reports that scientific idea generation is poorly predicted by standard metrics of general intelligence and that models with large gaps in general intelligence scores can nevertheless achieve comparable creative performance [2412.17596]. AGC-Bench similarly concludes that creativity is related to but separable from general knowledge and reasoning [2607.01152].

Third, scaling does not induce a uniform creativity gain. NoveltyBench reports that larger models within a family often exhibit *less diversity* than smaller counterparts [2504.05228]. CreativeBench describes this as “convergence-by-scaling”: larger models become more correct but less divergent, with scaling helping combinatorial creativity more than exploratory creativity [2603.11863]. C²-Eval likewise reports non-monotonicity with model size and larger gains from reasoning augmentation than from scale alone [2510.04009].

Fourth, prompting can help, but usually incompletely. NoveltyBench finds that paraphrasing prompts or telling a model to “be creative” has little impact, whereas in-context regeneration can greatly improve diversity, though this is externally forced rather than intrinsic to the single-prompt distribution [2504.05228]. AGC-Bench reports that prompting models to “be creative” yields a **1.40 SD** increase, much larger than enabling explicit reasoning (**0.34 SD**) [2607.01152]. CREATE finds only limited gains from simple creative prompting, with iterative regeneration providing the largest consistent boost in utility and distinctiveness [2603.09970].

Finally, human superiority persists in domains where taste, interpretive transfer, or real-world ambiguity matter. GuessBench reports that even GPT-4o is incorrect on **34\%** of instances and that open and API models differ dramatically in average performance (**13.87\%** vs. **53.93\%**) [2506.00814]. HCB finds that no model excels uniformly across ideation, mockup, and refinement [2606.30561]. The marketing Creativity Benchmark finds tightly clustered performance, with top-bottom spread \(\Delta\theta \approx 0.45\), implying a head-to-head win probability of only **0.61**, so that no model dominates across brands or prompt types [2509.09702].

## 6. Methodological debates and future directions

A persistent controversy is whether creativity can be summarized by a single score. Some benchmarks explicitly resist that reduction. CreativityPrism reports strong correlations between diversity and quality metrics but much weaker correlation of novelty with either, supporting the claim that strong performance in one creativity task or dimension does not necessarily generalize to others [2510.20091]. C²-Eval formalizes this through separate assessments of Usefulness, Originality, and Surprise across convergent and divergent regimes [2510.04009]. HCB argues more fundamentally that disagreement itself is informative: convergence concentrates on technical correctness and visual hierarchy, while divergence concentrates on aesthetic direction and conceptual risk [2606.30561].

Another debate concerns contamination, memorization, and the boundary between creativity and retrieval. CrIB addresses this by hiding targets and limiting the knowledge base [1805.03720]. CS\(^4\) raises prompt specificity to hinder retelling from the training corpus and reports that Learning from Human Feedback helps models select better stories from training data but has limited influence on producing creative stories unseen in the corpora [2410.04197]. LoTbench, built around the Oogiri game, addresses information leakage and interpretability through an interactive, causality-aware framework that scores how quickly a model converges toward high-quality human-level creative responses [2501.15147]. CreativeBench uses executable code and automated test cases to distinguish genuine novelty from non-functional hallucination [2603.11863].

Future work in the benchmark literature is oriented toward broader coverage, better human alignment, and more controllable creativity. AGC-Bench releases a public leaderboard, AGC-Judge, and human data as open infrastructure for large-scale measurement [2607.01152]. CreBench and CreMIT seek to train MLLMs that understand human-aligned creativity from idea to process to product [2511.13626]. CREward proposes a type-specific reward model for geometry, material, and texture, extending creativity assessment into image formation pipelines and explainable generation [2511.19995]. Across these efforts, a shared implication emerges: evaluation must distinguish where models should converge toward professional standards from where they should preserve pluralism, risk, and steerability. This suggests that the future of creativity benchmarking will be less about a universal scalar and more about structured measurement of multiple creative capacities under explicitly different operational definitions.

Source: https://www.emergentmind.com/topics/creativity-benchmark