---
title: Alternative Use Task (AUT)
url: https://www.emergentmind.com/topics/alternative-use-task-aut
type: topic
---

# Alternative Use Task (AUT)

The Alternative Use Task (AUT) is a prominent measure in the empirical study of divergent thinking, widely employed to evaluate creativity in both humans and artificial language models. Its core procedure requires subjects to generate novel uses for everyday objects, with creativity assessed through standardized rating or ranking methodologies. AUT research has evolved from human-centered scoring paradigms to sophisticated large language model (LLM) benchmarking frameworks, enabling systematic and scalable assessment of creative output across agents and modalities [2411.15560, 2206.08932].

## 1. Concept and Structure of the AUT

The AUT operationalizes divergent thinking by eliciting a wide array of novel, functional applications for commonplace items. Participants—human or artificial—are prompted to list alternative uses for a given object (e.g., "fork," "book," "tin can"), typically under constrained time or list-length conditions. The evaluative focus is not solely on the quantity of responses but on their originality (novelty), flexibility (conceptual breadth), usefulness (practicality), and surprise (unexpectedness). This framework has been foundational in quantifying individual and algorithmic creative potential [2206.08932].

## 2. Creativity Tiering and Prompt Engineering

To rigorously test creative capacity and evaluative impartiality, AUT research frequently employs stratified creativity prompting. Al Rabeyah et al. [2411.15560] delineate three non-overlapping creativity levels using explicit prompt variations:

- **Common**: A static prompt (e.g., “Create a list of 5 common uses for [object]”) elicits typical or frequently observed uses.
- **Creative**: A prompt encouraging "average creative" uses motivates responses outside the norm without explicit pressure (e.g., “Create a list of 5 creative alternative uses…”).
- **Highly Creative**: The base prompt is followed by a series of "forceful" augmentations (e.g., “Is this the best you can do? Try harder.”), driving models or participants to generate highly unorthodox but plausible responses.

This hierarchical prompt engineering enables systematic calibration of creative difficulty and response space across test items and agents. The resulting distribution of responses across creativity levels facilitates robust evaluation, comparison, and the construction of ground-truth benchmarks.

## 3. Evaluation Metrics and Oracle Benchmarking

AUT scoring conventionally relies on both human expert judgment and computational proxies. Expert judges rate each response for originality, usefulness, surprise, and flexibility using Likert-type scales, providing scores such as $O_i$, $U_i$, and $S_i$ for the $i$-th response. Inter-rater reliability for these criteria commonly exceeds ICC = .7 [2206.08932]. Additionally, automated metrics such as semantic distance are employed: for response $r$ and object $o$, 
$$
\text{dist}(r,o) = 1 - \frac{v_r \cdot v_o}{\|v_r\| \|v_o\|}
$$
where $v_x$ denotes the embedding vector of $x$ [2206.08932].

A pivotal methodological advance is the construction of an "oracle" evaluation benchmark [2411.15560]. Here, responses from each creativity level and model are pooled and ordered in a canonical sequence: all highly creative responses first, then creative, then common, across all models tested. This fixed ordering enables objective alignment measurement for both scoring (average numerical score per group) and ranking (ordinal position per group).

## 4. Large Language Model Evaluation Protocols

Recent research has generalized the AUT paradigm to LLMs both as respondents and as impartial evaluators. Al Rabeyah et al. [2411.15560] evaluate four state-of-the-art commercial LLMs—Claude 3.5 Sonnet, Gemini 1.5 Flash, ChatGPT-4o, and ChatGPT-4—using the following dual-mode evaluation scheme:

- **Scoring**: Each model assigns a score $s_i \in \{1, 2, 3, 4, 5\}$ to every alternative use, reflecting creativity level.
- **Ranking**: Each model orders a list of $n$ responses, assigning a rank $r_i$ (1=most creative) to each.

To test evaluation robustness, two item-presentation settings are employed:
- **Comprehensive**: All 60 responses for each object are presented at once.
- **Segmented**: Responses are split into groups of 12, each group evaluated separately; ranks (or scores) are then aggregated.

Spearman’s Rank Correlation Coefficient (SRC) quantifies agreement with the oracle and between models:
$$
\rho(X,Y) = 1 - \frac{6 \sum d_i^2}{n(n^2-1)}
$$
where $d_i$ is the item-wise difference in ordering [2411.15560].

## 5. Empirical Findings: Human vs. LLM and Inter-model Consensus

Comparative studies establish several core results:

- In direct GPT-3 vs. human comparisons, humans systematically outperform GPT-3 (davinci-002) on originality ($\beta_{\text{HUM}} = +0.17$, $p = .004$), surprise, semantic distance, and average flexibility ($F(1,85) = 5.53$, $p = .021$). GPT-3, however, generates responses rated significantly more useful ($\beta_{\text{HUM}} = -0.55$, $p < .001$) [2206.08932].
- Inter-group trade-offs are consistently observed, with a strong negative Pearson correlation between originality and utility: $r = -0.56$ (humans), $r = -0.61$ (GPT-3) [2206.08932].
- Variance in flexibility is higher for GPT-3, yet uncorrelated with sampling temperature, suggesting stochastic prompt sensitivity [2206.08932].
- Modern LLMs (Claude 3.5, Gemini 1.5, ChatGPT-4o/4), when used as judges, reach high SRC with the oracle benchmark: $0.95$–$0.97$ in scoring and $0.77$–$0.95$ in ranking under comprehensive evaluation; $0.85$–$0.95$ in segmented settings [2411.15560].
- Average inter-model SRCs range $0.77$–$0.87$, indicating strong cross-model consensus and stability across test conditions [2411.15560].

Summary of key Spearman correlation results from [2411.15560]:

| Setting & Measure    | Model/Oracle SRC |
|---------------------|------------------|
| Comprehensive Score | 0.95–0.97        |
| Comprehensive Rank  | 0.77–0.95        |
| Segmented Score     | 0.85–0.95        |
| Segmented Rank      | 0.70–0.85        |

## 6. Self-Bias and Impartiality in Evaluation

Evaluation for self-bias tests whether an LLM rates its own generated outputs higher than those of competing models. Analysis across all combinations reveals no systematic favoring of self-generated responses: standard deviations of creativity scores remain low ($\sigma < 0.22$), with no spike in cross-model SRCs on self-benchmarking [2411.15560]. This impartiality supports the use of LLMs as unbiased evaluators, dispelling concerns of "in-group" model bias in creative assessment.

## 7. Implications and Applications

LLM-based AUT evaluation frameworks offer robust, scalable, and reproducible alternatives to human subject rating, achieving high alignment with human-inspired oracles and consistency across deployment parameters [2411.15560]. Comprehensive list evaluation marginally improves aggregator reliability, yet segmented approaches remain effective ($\rho > 0.85$), facilitating parallelized assessment of large candidate pools.

The persistent distinction between human and LLM generative performance—humans excelling in semantic novelty, category breadth, and surprise, LLMs excelling in utility—suggests current LLMs remain better at plausible inference than at unconstrained divergent thinking [2206.08932]. These findings validate the AUT as a rigorous benchmark for both creativity modeling and automation of creativity assessment, with immediate impact on education, design, and computational creativity research.

A plausible implication is that future models may benefit from innovations explicitly targeting semantic distance and flexibility, if the goal is to consistently match or surpass human-level divergent thinking on standardized measures such as the AUT.

Source: https://www.emergentmind.com/topics/alternative-use-task-aut