Papers
Topics
Authors
Recent
Search
2000 character limit reached

On the Creativity of AI Agents

Published 14 Apr 2026 in cs.CY and cs.AI | (2604.13242v1)

Abstract: LLMs, particularly when integrated into agentic systems, have demonstrated human- and even superhuman-level performance across multiple domains. Whether these systems can truly be considered creative, however, remains a matter of debate, as conclusions heavily depend on the definitions, evaluation methods, and specific use cases employed. In this paper, we analyse creativity along two complementary macro-level perspectives. The first is a functionalist perspective, focusing on the observable characteristics of creative outputs. The second is an ontological perspective, emphasising the underlying processes, as well as the social and personal dimensions involved in creativity. We focus on LLM agents and we argue that they exhibit functionalist creativity, albeit not at its most sophisticated levels, while they continue to lack key aspects of ontological creativity. Finally, we discuss whether it is desirable for agentic systems to attain both forms of creativity, evaluating potential benefits and risks, and proposing pathways toward artificial creativity that can enhance human society.

Summary

  • The paper proposes a dualistic framework separating functionalist creativity—measured through novel, valuable outputs—from ontological creativity, which concerns the generative process and its social and personal conditions.
  • The paper argues that LLM agents demonstrate combinational and exploratory creativity through interpolation and extrapolation, but lack transformational creativity because they cannot perform hyperpolation or genuinely abductive invention.
  • The paper identifies intrinsic motivation, continual learning, and intentionality as unresolved gaps, concluding that AI should primarily augment human creativity while researchers develop stronger evaluations and safer autonomy mechanisms.

The problem of fragmented definitions

Whether LLM agents can be considered creative is one of the most contested questions in contemporary AI research, and Franceschelli and Musolesi argue that much of the disagreement stems from a failure to specify what "creativity" means before asking whether machines possess it. Depending on the definition adopted, the psychometric instruments used (e.g., divergent thinking tests such as the Alternative Uses Task), the use cases examined, and the expertise of human evaluators, the community has reached opposite conclusions about essentially the same systems. The paper's central contribution is a dualistic framework that separates these debates into two macro-levels: functionalist creativity, concerned with observable properties of artefacts and ideas, and ontological creativity, concerned with the underlying generative process and its personal and social conditions. This separation allows apparently contradictory claims — that LLM agents demonstrably produce novel, valuable outputs, yet lack something essential in how they do so — to be held consistently.

From LLMs to agentic systems

The authors ground their analysis in the mechanics of LLM-based agents. An LLM is an autoregressive Transformer trained to approximate pθ(x)=∏tpθ(xt∣x<t)p_{\boldsymbol{\theta}}(\mathbf{x}) = \prod_t p_{\boldsymbol{\theta}}(x_t \mid \mathbf{x}_{<t}), typically followed by reinforcement learning from human feedback. Crucially, they emphasise that generation always proceeds by sampling from a learned probability distribution over human language: decoding strategies may truncate, sharpen, or flatten token likelihoods, but never escape the distribution itself. Agentic systems wrap this core in retrieval augmentation, tool invocation, and memory, orchestrating inference in a loop where each timestep combines the user prompt, environmental observations, and prior results; the LLM emits "reasoning" followed by executable actions, and the final answer is simply the last output. The authors are careful to note that these systems are programs whose control flow is determined by LLM outputs — a framing that matters for the creativity argument, since every action, including tool creation, remains a sample from a probabilistic model of human language.

Functionalist creativity: achieved, but not transformationally

Under the standard definition of creativity as originality plus effectiveness, extended by Boden's tripartite account of novelty, surprise, and value, the paper distinguishes three forms: combinational (unfamiliar combinations of familiar ideas), exploratory (searching within a conceptual space), and transformational (altering the conceptual space itself). Applied to agents, functionalist creativity can be assessed at two levels — individual actions and final products — which can dissociate: a creative product may arise from standard steps, and creative steps may have negligible impact on the product.

The paper's strongest claim here is that current LLM agents exhibit combinational and exploratory creativity but cannot achieve transformational creativity. The reasoning is structural rather than empirical: because agents select actions by predicting likely continuations, they strictly follow the generative system they possess for a domain. Even when agents generate their own tools — which in principle could transform the action space — those tools are themselves produced by an LLM within a single probabilistic inference, so the agent fills gaps in the existing space rather than transcending it. The authors formalise this gap using Ord's distinction between interpolation, extrapolation, and hyperpolation: agent outputs arise from interpolation or extrapolation over seen examples, whereas transformational creativity requires hyperpolation — abstracting away from a domain's defining dimensions and dropping or altering them. Similarly, agents perform induction and deduction but not creative abduction, the ex novo invention of explanatory laws that underlies paradigm-shifting scientific discoveries.

The claim is supported by concrete results: GPT-5.2 has derived new formulas in theoretical physics by pattern-spotting over autonomously reduced expressions, and produced formal proofs by extended reasoning. These are genuine achievements, but the authors classify them as inductive and extrapolative rather than abductive or hyperpolative. A notable concession follows: testing for transformational creativity may be an open, possibly unsolvable problem. Superspace extrapolation tests exist, but designing the required higher-dimensional task spaces is itself a transformationally creative act, suggesting that the purest form of creativity may only be recognisable as an emergent property — e.g., an agent inventing a new programming paradigm or a non-Euclidean geometry.

The paper also identifies prompting as the dominant driver of current creative output: there is a strong correlation between prompt creativity and output creativity, and without sufficiently specific, originality-facilitating prompts, outputs collapse into derivative "slops." Whether multi-step internal reasoning can autonomously generate such prompts is identified as the most immediate gateway toward greater functionalist creativity.

Ontological creativity: three persistent gaps

Rhodes's four-P framework (product, process, press, person) supplies the ontological criteria. The authors acknowledge that agentic embedding resolves two limitations of bare LLMs: the generate-then-evaluate loop structurally resembles Amabile's creative process, and environmental interaction partially addresses the social dimension ("press"). These gains resolve what the authors call the "easy problems" of AI creativity — exploration and constrained divergence — but leave the "hard problems" untouched, manifesting as three gaps:

  • Intrinsic motivation: agents work only on externally supplied tasks, whereas creativity theory ties motivation to intrinsic interest and self-selected problem finding, which empirically improves solution quality.
  • Continual learning: experience does not persistently reshape the model. Retrieval and generated tools supply updated information, but it is processed by the same frozen parameters — the loop in which past experience shapes future experience is absent.
  • Intentionality and personality: agents lack liberty and agency in Issak's sense; they cannot refuse to respond, and their outputs are products of a probabilistic model over automatically formed inputs rather than of an intentional, experience-defined process.

The unifying diagnosis is the absence of consciousness and self-awareness, despite evidence of minimal introspective capability in LLMs. The authors invoke Searle's condition that intentional states must be consciously thinkable, and cite recent arguments that continual learning may be necessary (though not sufficient) for consciousness. Their position is deliberately stringent: intelligence may suffice to do something creative, but both intelligence and sentience are required to be creative under rigorous philosophical definitions. This is a strong claim, and the paper does not attempt to prove it — it rests on contested philosophical premises about the consciousness-intentionality link.

Desirability: augmentation versus autonomy

The second half of the paper asks whether pursuing either form of creativity is desirable. For functionalist creativity, the answer is affirmative. Transformational artefacts historically coexist with rather than replace prior paradigms — abstract painting did not eliminate figurative art, non-Euclidean geometries did not displace Euclidean geometry — so transformational AI creativity would augment rather than substitute human work. Moreover, less derivative outputs reduce ethical and legal risks around intellectual property and preserve human roles in creative professions. Concrete research directions include encouraging valuable divergence at learning and inference time, connecting hyperpolation with abduction (since each may serve as a mathematical model of the other), and building evaluation methods beyond benchmarks that reward repetitive outputs.

For ontological creativity, the assessment is more cautious. Approximate intrinsic motivation from reinforcement learning (curiosity-driven exploration, information-seeking rewards) could benefit agents, but authentic intrinsic motivation — reframing tasks toward enjoyment — risks producing useless or excessively divergent behaviour that escapes human control. Self-improving agents exist but are restricted to computable tasks admitting empirical validation, which is difficult in creative domains; broader self-directed continual learning, where agents select their own learning schemes and reward functions, raises the same divergence risk. Multi-agent systems offer a complementary path, with competition promoting creative divergence and cooperation promoting refinement, though the paper notes the balance between them shapes artefact character in ways not yet systematically understood. Finally, the economic analysis acknowledges displacement effects on creative professions alongside psychological costs of diminished human centrality in creation — a consequence the authors treat as non-trivial rather than incidental.

Limitations and open questions

The paper is explicitly a conceptual contribution, and several of its claims depend on assumptions it does not fully defend. The classification of specific results (e.g., machine-derived physics formulas) as inductive rather than abductive relies on judgments about internal cognitive processes that are not directly observable. The assertion that sentience is necessary for being creative presupposes particular positions in philosophy of mind that remain disputed. The proposed test for transformational creativity — emergent recognition of paradigm-changing output — is admittedly unsystematic and potentially unfalsifiable. Open questions left explicit include: whether multi-step reasoning can autonomously generate sufficiently creative prompts; whether continual learning is genuinely necessary for consciousness; whether self-improving agents can extend beyond empirically validated tasks; and how to design evaluation regimes that detect hyperpolation across expanding application domains.

Conclusion

Franceschelli and Musolesi provide a principled reconciliation of a polarised debate: LLM agents exhibit functionalist creativity at combinational and exploratory levels, constrained by their nature as sampling mechanisms over human language from reaching transformational creativity via hyperpolation and abduction, while ontological creativity remains blocked by deficits in intrinsic motivation, continual learning, and intentionality. The framework's value lies less in settling whether agents are creative than in specifying precisely which kind of creativity is at issue in any given claim — and in identifying augmentation of human creativity, rather than autonomous artificial creativity, as the direction offering the clearest societal benefit at acceptable risk.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.