---
title: 'PixelHumor: Multimodal Comic Benchmark'
url: https://www.emergentmind.com/topics/pixelhumor
type: topic
---

# PixelHumor: Multimodal Comic Benchmark

Searching arXiv for PixelHumor and closely related multimodal humor benchmarks.
PixelHumor is a benchmark dataset of **2,800 annotated multi-panel comics** designed to evaluate **Large Multimodal Models’ (LMMs) ability to interpret multimodal humor and recognize narrative sequences**. It targets online comics rather than single-panel or unimodal material, and therefore places emphasis on **multimodal contextual and narrative reasoning**, including punchline localization, modality attribution, humor-style recognition, open-ended interpretation, and sequence reconstruction. In reported experiments, current LMMs perform strongly on some surface-level tasks yet remain substantially below human performance on narrative and humor comprehension, especially when humor depends on cross-panel buildup or fine-grained integration of visual and textual cues [2509.12248].

## 1. Corpus design and scope

PixelHumor is built from **2,800 multi-panel online comics** drawn from seven webcomic sources: **Cyanide and Happiness**, **Peanuts**, **Garfield**, **XKCD**, **PhD Comics**, **They Can Talk**, and **Saturday Morning Breakfast Cereal (SMBC)** [2509.12248]. The corpus is explicitly intended to cover both **textual and visual humor** across “diverse creative and stylistic traditions,” which distinguishes it from humor resources centered on single images, captions, or text-only jokes.

The dataset statistics reflect substantial structural variation. Comics contain **up to 18 panels**, with the average varying by source, and the number of words ranges from **0** in silent comics to **501**. Reported source-level averages span from **~2.5 panels for XKCD** to **4.5 for Peanuts**. This heterogeneity is significant because it forces models to handle not only local image-text associations but also differences in strip length, dialogue density, and narrative pacing [2509.12248].

PixelHumor is also framed as a copyright-conscious benchmark. Its release is described as **“links with metadata for copyright compliance,”** which positions it as an evaluation framework rather than a conventional fully redistributed image corpus [2509.12248]. This suggests a design optimized for reproducible benchmarking while respecting source constraints.

## 2. Annotation protocol and representational schema

The annotations were produced by **8 trained undergraduates (18–25 y/o)**, with each comic **double-annotated for reliability** and a **3rd annotator adjudication** step when disagreements greater than **15%** arose [2509.12248]. PixelHumor provides both **per-example aggregated labels** and **individual annotator responses**, which is unusually useful for work on disagreement modeling, calibration, and subjectivity in humor research.

The annotation scheme has five components:

1. **Humor presence**: whether the comic is intended as humorous.  
2. **Sound effects**: whether sound effects such as onomatopoeia contribute to humor.  
3. **Critical panel detection**: which panel delivers the punchline.  
4. **Modality attribution**: whether humor arises mainly from text, image, or both.  
5. **Humor style labeling**: multi-label assignment using an **8-style taxonomy** [2509.12248].

The humor-style taxonomy comprises **Comparison**, **Personification**, **Exaggeration**, **Pun**, **Sarcasm**, **Silliness**, **Surprise**, and **Dark**. Reported style frequencies are uneven: **Surprise** is the most common at **35%**, followed by **Personification** at **28%**, while **Dark** is least frequent at **5%** and is concentrated in particular sources [2509.12248]. The modality distribution is similarly nonuniform. **Text-driven humor** accounts for **52%** of comics, **joint visual-text** humor for **32%**, and **sound effects** are present in **~15%** of comics, with **70%+** of those cases directly enhancing the humor [2509.12248].

Annotation quality is summarized with two different agreement measures. **Overall agreement** is reported as **0.872 (exact matches/multi-label overlap)**, while **Krippendorff’s alpha** is lower at **0.556**, particularly for subjective labels such as **“Do you get the joke?”** The paper attributes this lower alpha largely to the **strong humor-skew in source material** [2509.12248]. A common misconception is that a high overlap score implies that humor annotation is nearly solved; the PixelHumor statistics instead indicate that subjective humor judgments can remain difficult even under trained annotation protocols.

## 3. Benchmark tasks and evaluation methodology

PixelHumor defines **four core tasks**, each targeting a different layer of multimodal humor understanding [2509.12248].

| Task | Objective | Metrics |
|---|---|---|
| Humor Identification Tests | Humor presence, sound effect detection, critical panel detection, modality attribution | Precision, recall, F1-score |
| Humor Classification | Multi-label humor style classification | Micro/macro precision, recall, F1 |
| Humor Interpretation | Explain why the comic is or is not funny in 3 sentences | 7-point Likert ratings for relevance and coherence |
| Sequence Recognition | Reconstruct correct panel or dialogue order | Panel sequence accuracy, text sequence accuracy, WER, CER |

The **Humor Identification Tests** are discriminative subtasks that probe whether an LMM can detect coarse properties of a comic. These are not limited to humor detection; they also test whether a model can identify the **critical panel** and infer whether humor is primarily **textual**, **visual**, or **joint** [2509.12248].

The **Humor Classification** task is explicitly **multi-label**. This matters because a comic may simultaneously instantiate, for example, **Surprise** and **Sarcasm**, and the benchmark therefore penalizes models that collapse multi-style humor into a single dominant label [2509.12248].

The **Humor Interpretation** task requires a natural-language explanation in **3 sentences**. It is evaluated by humans using **7-point Likert scale ratings** for **relevance** and **coherence**, rather than by lexical overlap alone. This is consistent with the benchmark’s view that understanding humor is not reducible to label prediction [2509.12248].

The **Sequence Recognition** task directly targets narrative competence. It measures whether a model can reconstruct the reading order of **panels** or **dialogue**, using **panel sequence accuracy**, **text sequence accuracy**, **Word Error Rate (WER)**, and **Character Error Rate (CER)**. The reported WER definition is:

$$
\mathrm{WER} = \frac{S + D + I}{N}
$$

where \(S\) is substitutions, \(D\) is deletions, \(I\) is insertions, and \(N\) is the number of words in the reference [2509.12248].

## 4. Experimental results with large multimodal models

PixelHumor evaluates both **closed-source** and **open-source** LMMs in a **zero-shot** setting with **standardized prompts**. The closed-source models are **GPT-4o** and **Gemini-1.5-Pro**; the open-source set includes **Qwen2-VL-72B**, **Gemma3-27B**, **LLaVA-OneVision-7B**, and **Qwen2-VL-7B** [2509.12248].

The headline pattern is asymmetric. On **Humor Presence**, models reach **F1 ≈ 0.98**, which is near ceiling. However, the same experiments report that this performance reflects **dataset bias**, since most comics are humorous and models have **nearly random performance distinguishing genuinely non-humorous ones** [2509.12248]. By contrast, the more diagnostic tasks remain difficult.

On **Sound Effect Detection**, **GPT-4o** achieves **F1: 0.82**, while typical open models score **0.70–0.74**. On **Panel Contribution**, the best reported result is **GPT-4o with F1: 0.77**, whereas open models are in the **0.49–0.54** range. On **Modality Attribution**, **GPT-4o** reaches **F1: 0.63**, while open models span **0.21–0.58** [2509.12248].

The **Humor Style** task is harder still. The reported best **total F1** is **0.50** for **GPT-4o**, while open models range from **0.09 (LLaVA)** to **0.48 (Gemini)** in the synopsis table [2509.12248]. The qualitative account states that models perform best on **highly visual/personification jokes**, while **sarcasm**, **dark**, and **subtle humor types** are poorly detected.

Narrative ordering remains a major bottleneck. On **Sequence Recognition**, the best listed model is **Gemini-1.5-Pro** with **Panel Acc: 64.5%**, whereas typical open LMMs achieve **31–42%**, and the measured **human baseline is 100% on the tested subset** [2509.12248]. This gap is central to PixelHumor’s scientific value: it shows that high-capacity multimodal systems still struggle with the sequential structure that underpins comic timing.

For **Humor Interpretation**, **GPT-4o** obtains a mean of **5.8/7**, while open models fall in the **3.0–4.6/7** range. Yet even here, **human-written explanations are consistently preferred**: in **69% of rating cases**, humans selected the human explanation over any model output [2509.12248]. This indicates that current LMM explanations often sound plausible without capturing the specific mechanism of the joke.

## 5. Failure modes, diagnostic value, and common misconceptions

PixelHumor’s central empirical conclusion is that **surface-level recognition is nearly solved**, whereas **true comprehension is largely unsolved** [2509.12248]. The benchmark therefore separates two questions that are often conflated: whether a model can identify that a comic is intended as humor, and whether it can explain why the humor works.

Several recurrent failure modes are reported. One is **heuristic bias**: models default to **“both modalities”** for modality attribution or assume a canonical **left-to-right reading order** even when the panel order has been randomized. Another is **hallucination**, in which models generate plausible but factually incorrect explanations. A third is reliance on **template explanations**, such as generic statements about “unexpected twist” or “absurd situation,” without pinpointing the relevant narrative-visual interaction [2509.12248].

Performance also degrades with comic length. The benchmark reports that **longer comics** produce **weaker interpretation performance**, which indicates difficulty with **extended narrative/factual tracking** across multiple panels [2509.12248]. This is consistent with the broader implication drawn in the paper: LMMs need better **hierarchical/context tracking** to maintain comedic buildup and punchline delivery.

PixelHumor also highlights a misconception about multimodal scaling. Although **GPT-4o** and **Gemini** consistently outperform open models on nontrivial tasks, the benchmark concludes that increasing general multimodal capability has not closed the gap on **contextual integration**, **cultural and social nuance**, **narrative reasoning**, and **human alignment** [2509.12248]. The benchmark’s significance lies precisely in making these deficits measurable.

## 6. Position within computational humor research

PixelHumor occupies a specific niche within a larger humor-research landscape. It differs from **HumorDB**, which is an **image-only** dataset of **3,545 images** and **1,271 pairs** designed for **binary classification**, **range regression**, and **pairwise comparison** on graphical humor; HumorDB emphasizes **subtle visual cues** and paired contrasts, whereas PixelHumor centers **multi-panel narrative** and **multimodal comic reasoning** [2406.13564]. It also differs from **OxfordTVG-HIC**, which provides **approximately 2.9M image-text pairs with humour scores** to support humorous caption generation, including a **position-conditioned loss** for captioning models [2307.11636], and from the earlier **Neural Joking Machine**, which used **ResNet-152 + LSTM**, **Funny Score**, and **BoketeDB** for humorous image captioning [1805.11850].

In video, adjacent benchmarks and datasets address complementary phenomena. **ExFunTube** contains **10,136 user-generated short-form funny videos** with **timestamps and detailed human-written explanations**, emphasizing humor arising from the interaction of **visual and verbal cues** [2310.14159]. **CVLA** targets **short-form video humor detection** through **comment-aided video-language alignment** and multi-modal contrastive pre-training on **DY11k** and **UR-FUNNY** [2402.09055]. **V-HUB** is a **visual-centric humor understanding benchmark** for video LLMs built from **960 videos**, mostly with **visual-centric humor**, and shows that all tested models suffer a clear degradation when moving from text-based to raw video evaluation [2509.25773].

PixelHumor is also related to recent work on internet memes and online communities. The study of **r/ProgrammerHumor** reports that **image-based submissions are significantly more likely to achieve higher Reddit scores than text-only submissions**, with **80%** of annotated images adding context to the humor, which reinforces the broader relevance of multimodal humor analysis beyond comics [2410.07020]. At the generation end, the **HUMOR** framework treats meme creation as a multimodal problem requiring **hierarchical, multi-path Chain-of-Thought**, **group-wise pairwise reward modeling**, and **group-wise reinforcement learning** aligned to within-template human preferences [2512.24555].

Within this landscape, PixelHumor’s distinctive contribution is to make **comic-specific multimodal humor understanding** measurable at the intersection of **narrative sequencing**, **style recognition**, **punchline localization**, and **explanatory reasoning**. This suggests that PixelHumor is not merely another humor dataset, but a benchmark for probing whether LMMs can integrate visual and textual evidence across time-like panel structure in a way that approximates human comic comprehension [2509.12248].

Source: https://www.emergentmind.com/topics/pixelhumor