Papers
Topics
Authors
Recent
Search
2000 character limit reached

Oogiri-Master Humor Benchmark

Updated 16 July 2026
  • Oogiri-Master is a benchmark for humor understanding that uses the Japanese creative game Oogiri to evaluate LLMs.
  • It leverages the Oogiri-Corpus with 908 prompts and over 82,000 responses, ensuring robust evaluation through independent, hidden human votes.
  • The benchmark demonstrates that insight-augmented prompting and detailed feature analysis can significantly enhance model performance.

Oogiri-Master is a benchmark for evaluating humor understanding in LLMs through the Japanese creative response game Oogiri, in which participants produce witty responses to a prompt. It is paired with Oogiri-Corpus, a dataset designed to support rigorous analysis of what makes a response funny to humans. The benchmark addresses three limitations identified in earlier work: datasets with few candidate responses per prompt, exposure of popularity signals during rating, and the absence of objective and comparable metrics for funniness. Its central contribution is to combine large candidate sets per prompt, independently collected human judgments, quantitative feature analysis, and benchmark tasks that test relative and absolute humor judgments in a controlled way (Murakami et al., 25 Dec 2025).

1. Conceptual scope and research objective

Oogiri-Master treats humor as a testbed for human-like creative thinking in LLMs. Its framing is narrower than generic joke generation: the focus is not only whether a model can produce or recognize something “funny,” but whether it can evaluate funniness under the prompt-conditioned, open-ended constraints of Oogiri. In this game format, prompts are short and creative, and success depends on producing or identifying responses that are witty relative to the specific prompt rather than merely unusual in isolation (Murakami et al., 25 Dec 2025).

The benchmark is built around a concrete research question: what makes Oogiri responses funny to humans. The work explicitly argues that prior resources offered only limited reliable means to answer this question, because they contained few candidate responses per prompt, revealed popularity signals during rating, and lacked objective and comparable metrics for funniness. Oogiri-Master and Oogiri-Corpus were introduced to enable rigorous evaluation under those constraints.

A useful distinction within the Oogiri literature is between benchmarks centered on generation and those centered on evaluation. Oogiri-Master primarily operationalizes humor understanding through judgment tasks, whereas related work has also used Oogiri for generation and creative reasoning. This suggests that Oogiri-Master occupies a specific methodological niche: it is a benchmark for evaluating how closely model judgments track aggregated human funniness judgments, rather than a general-purpose platform for all humor tasks.

2. Oogiri-Corpus: dataset structure and human annotation

Oogiri-Corpus consists of short, open-ended, creative prompts paired with many independently written responses. Each prompt is paired with approximately 100 diverse candidate responses; the reported mean is approximately 96 responses per prompt. The corpus contains 908 prompts and 82,536 responses, with an average response length of 16.4 characters and an average prompt length of 20.4 characters. It is described as about 7x the size of comparable datasets such as Oogiri-GO (Murakami et al., 25 Dec 2025).

Item Value Note
Prompts 908 Included in the corpus
Responses 82,536 Approximately 96 per prompt
Average prompt length 20.4 characters Short creative stimuli
Average response length 16.4 characters Witty replies
Average votes per prompt ≈ 172 ≈ 1.8 votes per response

The annotation protocol is designed to mitigate popularity bias. After the response phase, a voting phase is conducted in which approximately 100 different human judges independently vote on which responses they find funny. During voting, the number of votes per response is hidden, so raters do not see others’ judgments. The paper identifies this explicitly as a mechanism for preventing the “bandwagon effect.” The large number of responses per prompt also changes the evaluation regime: judges are not forced to choose the “least bad” option from a narrow set, but instead evaluate from a relatively diverse candidate pool.

Reliability is further controlled by filtering on vote count. On average, each prompt receives approximately 172 votes, corresponding to approximately 1.8 votes per response, and only prompts with at least 100 total votes are included in the benchmark. The stated purpose is to ensure evaluation reliability and reduce noise. In this design, robustness comes not from a single authoritative rating, but from aggregation over many independent judgments.

3. Quantitative decomposition of funniness

To analyze which linguistic factors are associated with funniness, the study selects the top 3 high-rated and bottom 3 low-rated responses for each prompt, using vote counts, and balances the sample for statistical comparison. The resulting analysis set comprises 5,448 analyzed pairs. This construction makes the feature analysis explicitly contrastive: the target is not humor in the abstract, but differences between responses that humans consistently preferred and those they consistently rejected (Murakami et al., 25 Dec 2025).

The analyzed features span several levels.

Basic features include sentence or response length, unique character count, character-type ratios such as hiragana and katakana, part-of-speech ratios, prompt–response length ratios, and lexical novelty defined as the fraction of unique words in the response not in the prompt.

Semantic features include semantic distance, defined as $1 -$ cosine similarity between prompt and response embeddings, and textual entailment probabilities for entailment, contradiction, and neutrality from NLI models.

Information-theoretic features include surprisal, defined as mean negative log likelihood per token under a LLM, and normalized pointwise mutual information, used to quantify unexpectedness of co-occurrence.

Higher-order features are scored by LLMs through custom prompts on a 1–5 scale. These features are ambiguity exploitation, associative distance, benign violation, coherence, expectedness, incongruity resolution, metaphor use, and perspective shift.

Differences between high-rated and low-rated responses are assessed using a two-sample Student’s tt-test, and effect size is measured with Cohen’s dd:

d=Xˉ1−Xˉ2sp,sp=(n1−1)s12+(n2−1)s22n1+n2−2d = \frac{\bar{X}_1 - \bar{X}_2}{s_p}, \quad s_p = \sqrt{\frac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1 + n_2 - 2}}

where Xˉ\bar{X} is the group mean, ss is the group standard deviation, nn is the group size, and sps_p is the pooled standard deviation.

The reported findings are specific. High-rated responses tend to be shorter and less lexically novel, which is interpreted as being less off-topic. Higher-order features, notably ambiguity, perspective shift, and incongruity resolution, are notably higher in funny responses and show moderate effect sizes. By contrast, other features such as part-of-speech ratios and surprisal are statistically significant but have small effect sizes. The overall picture is that prompt-conditioned humor in Oogiri is not reducible to raw unexpectedness alone; higher-order structure matters.

4. Benchmark task design and evaluation protocol

Oogiri-Master defines multiple task formats for humor judgment. The relative judgment tasks use a multiple-choice question answering format with binary, triple, and quadruple choice settings, in which the model must select the funniest response for a given prompt. The binary setting is split into two negative-sampling conditions: negatives drawn from the same prompt, which tests within-context funniness discrimination, and negatives drawn from a different prompt, which tests whether the chosen response is relevant to the prompt rather than merely generally funny (Murakami et al., 25 Dec 2025).

The benchmark also includes an absolute judgment task formulated as binary classification, where the model must decide whether a given response is “funny” or “not funny” for the prompt. Across all tasks, the total number of items is 600. These items are sampled to avoid overlap and contamination with the analysis set used for quantitative feature study, thereby separating descriptive analysis from held-out evaluation.

The model evaluation protocol is deterministic: temperature is set to zero. The tested systems include multiple proprietary models, such as GPT-5, Claude-Opus-4, and Gemini, and open-source models such as LLM-jp and DeepSeek-R1. Two prompting regimes are compared. Baseline prompts directly instruct the model to select the funniest response without additional guidance. Insight-augmented prompts provide additional context derived from the quantitative analysis, including features such as length and ambiguity.

The human baseline is constructed from Japanese crowdworkers who solve the same benchmark tasks under identical instructions. Each item is judged by 21 workers, and the majority vote is used as the gold label. Accuracy is the main metric and is averaged across tasks. This makes the reported model numbers directly comparable to an aggregated human decision rule rather than to a single annotator.

5. Reported performance and prompting effects

The headline result is that state-of-the-art models approach human performance on Oogiri-Master. Human performance is reported as 68.7% average accuracy across all tasks. Among baseline LLMs, Claude-Opus-4 also achieves 68.7%, and GPT-5 achieves 67.6%. Other models, including Gemini and the open-source systems, perform worse under the same evaluation protocol (Murakami et al., 25 Dec 2025).

System Average accuracy across all tasks Setting
Humans 68.7% Majority-vote baseline
Claude-Opus-4 68.7% Baseline prompt
GPT-5 67.6% Baseline prompt
GPT-5 70.7% Insight-augmented prompt

A second central result is that insight-augmented prompting can improve performance. The clearest reported example is GPT-5, which improves from 67.6% to 70.7%, surpassing both the human baseline and the baseline LLM results. The benchmark therefore does not merely measure ability; it also exposes a link between explanatory feature analysis and practical prompt design.

The gains, however, are not uniform across model classes. The paper reports that open-source or weaker models do not always benefit from insight augmentation, and that performance can sometimes drop because of over-reliance on shallow heuristics. Instruction style also matters: telling the model to “only use insights when uncertain” yields the biggest boost, which the authors attribute to reduced dependence on simple features.

The ablation findings reinforce this interpretation. Both basic features and higher-order features contribute, and simple features such as response length already yield substantial gains. Combined use of both feature types further improves the strongest models. The benchmark also reports that models with continued pretraining on Japanese improve, with DeepSeek-R1-Japanese presented as an example. The stated conclusion is that cultural and language adaptation is important for humor understanding.

6. Relation to adjacent Oogiri benchmarks and broader humor research

Oogiri-Master sits within a rapidly developing line of Oogiri-based humor research, but its methodological emphasis differs from neighboring efforts. One related study expands Oogiri data with new sources and annotates responses with six dimensions—Novelty, Clarity, Relevance, Intelligence, Empathy, and Overall Funniness—using 5-point absolute ratings from four native Japanese speakers (Sakabe et al., 12 Nov 2025). That study evaluates both generation and evaluation, and reports that state-of-the-art LLMs generate responses between low- and mid-tier human performance while showing a notable lack of Empathy. It also finds only modest human–model correlations on Overall Funniness and argues that humans prioritize Empathy whereas LLMs prioritize Novelty. This provides an important interpretive context for Oogiri-Master: near-human benchmark accuracy on selected judgment tasks does not necessarily imply human-like evaluative criteria.

Another line of work uses Oogiri to study creative reasoning rather than benchmarked humor judgment. “Let’s Think Outside the Box: Exploring Leap-of-Thought in LLMs with Creative Humor Generation” introduces the multimodal and multilingual Oogiri-GO dataset with over 130,000 samples and proposes the CLoT framework to improve Leap-of-Thought abilities in humor generation and discrimination (Zhong et al., 2023). That paper contrasts Chain-of-Thought with non-sequential associative reasoning and reports quantitative gains from CLoT across image-to-text, text-to-text, and multimodal Oogiri tasks. Relative to that agenda, Oogiri-Master is narrower and more controlled: it emphasizes rigorous evaluation of humor understanding, large independent human voting, and objective metrics associated with funniness.

Taken together, these studies indicate that Oogiri has become a productive benchmark family for at least three related but distinct problems: creative humor generation, multidimensional humor evaluation, and prompt-conditioned humor understanding. A plausible implication is that progress on one of these problems does not automatically transfer to the others. Oogiri-Master contributes most directly to the third of these by providing a principled benchmark tied to large-scale human judgments and feature-based analysis.

7. Significance, limitations, and interpretive cautions

The primary significance of Oogiri-Master is methodological. It provides objective, reliable funniness assessment through large-scale, independent human voting without popularity bias, and it links that assessment to a quantitative decomposition of linguistic factors such as text length, ambiguity, and incongruity resolution (Murakami et al., 25 Dec 2025). The benchmark therefore functions both as an evaluation suite and as an empirical framework for testing hypotheses about humor.

Its results also bear on the status of current LLMs. The paper demonstrates that state-of-the-art models can approach human performance, and that feature-aware prompting can even surpass the reported human crowdworker baseline on the benchmark. At the same time, the benchmark’s own ablations indicate that some of this improvement can come from simple cues such as response length, and related multidimensional work shows persistent divergence between human and model humor criteria, especially around empathy (Sakabe et al., 12 Nov 2025). These findings argue against equating benchmark success with full human-like humor competence.

A further caution concerns scope. Oogiri-Master is grounded in Japanese Oogiri and evaluates prompt-conditioned response judgments. Its conclusions are therefore strongest for this domain and task design. The paper itself presents the benchmark as a principled basis for evaluating and advancing humor understanding in LLMs, and as a benchmark for broader creative and cultural understanding. This suggests extensibility, but any generalization beyond Japanese Oogiri remains an inference rather than a direct empirical claim from the reported benchmark.

Within those bounds, Oogiri-Master establishes a technically precise reference point for computational humor research: large response sets per prompt, hidden popularity during rating, independently aggregated human votes, explicitly defined linguistic features, held-out judgment tasks, and direct comparison between human and model accuracy. Its importance lies less in proposing a single theory of humor than in supplying a rigorous experimental substrate on which multiple theories can be tested.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Oogiri-Master.