Papers
Topics
Authors
Recent
Search
2000 character limit reached

PixelHumor: Multimodal Comic Benchmark

Updated 10 July 2026
  • PixelHumor is a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate large multimodal models on narrative reasoning and humor comprehension.
  • It targets complex tasks including punchline localization, modality attribution, and humor style recognition to gauge models' multimodal contextual understanding.
  • Experiments show that while models perform well on surface-level humor detection, they struggle with narrative sequencing and fine-grained humor integration.

Searching arXiv for PixelHumor and closely related multimodal humor benchmarks. PixelHumor is a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate Large Multimodal Models’ (LMMs) ability to interpret multimodal humor and recognize narrative sequences. It targets online comics rather than single-panel or unimodal material, and therefore places emphasis on multimodal contextual and narrative reasoning, including punchline localization, modality attribution, humor-style recognition, open-ended interpretation, and sequence reconstruction. In reported experiments, current LMMs perform strongly on some surface-level tasks yet remain substantially below human performance on narrative and humor comprehension, especially when humor depends on cross-panel buildup or fine-grained integration of visual and textual cues (Ryan et al., 12 Sep 2025).

1. Corpus design and scope

PixelHumor is built from 2,800 multi-panel online comics drawn from seven webcomic sources: Cyanide and Happiness, Peanuts, Garfield, XKCD, PhD Comics, They Can Talk, and Saturday Morning Breakfast Cereal (SMBC) (Ryan et al., 12 Sep 2025). The corpus is explicitly intended to cover both textual and visual humor across “diverse creative and stylistic traditions,” which distinguishes it from humor resources centered on single images, captions, or text-only jokes.

The dataset statistics reflect substantial structural variation. Comics contain up to 18 panels, with the average varying by source, and the number of words ranges from 0 in silent comics to 501. Reported source-level averages span from ~2.5 panels for XKCD to 4.5 for Peanuts. This heterogeneity is significant because it forces models to handle not only local image-text associations but also differences in strip length, dialogue density, and narrative pacing (Ryan et al., 12 Sep 2025).

PixelHumor is also framed as a copyright-conscious benchmark. Its release is described as “links with metadata for copyright compliance,” which positions it as an evaluation framework rather than a conventional fully redistributed image corpus (Ryan et al., 12 Sep 2025). This suggests a design optimized for reproducible benchmarking while respecting source constraints.

2. Annotation protocol and representational schema

The annotations were produced by 8 trained undergraduates (18–25 y/o), with each comic double-annotated for reliability and a 3rd annotator adjudication step when disagreements greater than 15% arose (Ryan et al., 12 Sep 2025). PixelHumor provides both per-example aggregated labels and individual annotator responses, which is unusually useful for work on disagreement modeling, calibration, and subjectivity in humor research.

The annotation scheme has five components:

  1. Humor presence: whether the comic is intended as humorous.
  2. Sound effects: whether sound effects such as onomatopoeia contribute to humor.
  3. Critical panel detection: which panel delivers the punchline.
  4. Modality attribution: whether humor arises mainly from text, image, or both.
  5. Humor style labeling: multi-label assignment using an 8-style taxonomy (Ryan et al., 12 Sep 2025).

The humor-style taxonomy comprises Comparison, Personification, Exaggeration, Pun, Sarcasm, Silliness, Surprise, and Dark. Reported style frequencies are uneven: Surprise is the most common at 35%, followed by Personification at 28%, while Dark is least frequent at 5% and is concentrated in particular sources (Ryan et al., 12 Sep 2025). The modality distribution is similarly nonuniform. Text-driven humor accounts for 52% of comics, joint visual-text humor for 32%, and sound effects are present in ~15% of comics, with 70%+ of those cases directly enhancing the humor (Ryan et al., 12 Sep 2025).

Annotation quality is summarized with two different agreement measures. Overall agreement is reported as 0.872 (exact matches/multi-label overlap), while Krippendorff’s alpha is lower at 0.556, particularly for subjective labels such as “Do you get the joke?” The paper attributes this lower alpha largely to the strong humor-skew in source material (Ryan et al., 12 Sep 2025). A common misconception is that a high overlap score implies that humor annotation is nearly solved; the PixelHumor statistics instead indicate that subjective humor judgments can remain difficult even under trained annotation protocols.

3. Benchmark tasks and evaluation methodology

PixelHumor defines four core tasks, each targeting a different layer of multimodal humor understanding (Ryan et al., 12 Sep 2025).

Task Objective Metrics
Humor Identification Tests Humor presence, sound effect detection, critical panel detection, modality attribution Precision, recall, F1-score
Humor Classification Multi-label humor style classification Micro/macro precision, recall, F1
Humor Interpretation Explain why the comic is or is not funny in 3 sentences 7-point Likert ratings for relevance and coherence
Sequence Recognition Reconstruct correct panel or dialogue order Panel sequence accuracy, text sequence accuracy, WER, CER

The Humor Identification Tests are discriminative subtasks that probe whether an LMM can detect coarse properties of a comic. These are not limited to humor detection; they also test whether a model can identify the critical panel and infer whether humor is primarily textual, visual, or joint (Ryan et al., 12 Sep 2025).

The Humor Classification task is explicitly multi-label. This matters because a comic may simultaneously instantiate, for example, Surprise and Sarcasm, and the benchmark therefore penalizes models that collapse multi-style humor into a single dominant label (Ryan et al., 12 Sep 2025).

The Humor Interpretation task requires a natural-language explanation in 3 sentences. It is evaluated by humans using 7-point Likert scale ratings for relevance and coherence, rather than by lexical overlap alone. This is consistent with the benchmark’s view that understanding humor is not reducible to label prediction (Ryan et al., 12 Sep 2025).

The Sequence Recognition task directly targets narrative competence. It measures whether a model can reconstruct the reading order of panels or dialogue, using panel sequence accuracy, text sequence accuracy, Word Error Rate (WER), and Character Error Rate (CER). The reported WER definition is:

WER=S+D+IN\mathrm{WER} = \frac{S + D + I}{N}

where SS is substitutions, DD is deletions, II is insertions, and NN is the number of words in the reference (Ryan et al., 12 Sep 2025).

4. Experimental results with large multimodal models

PixelHumor evaluates both closed-source and open-source LMMs in a zero-shot setting with standardized prompts. The closed-source models are GPT-4o and Gemini-1.5-Pro; the open-source set includes Qwen2-VL-72B, Gemma3-27B, LLaVA-OneVision-7B, and Qwen2-VL-7B (Ryan et al., 12 Sep 2025).

The headline pattern is asymmetric. On Humor Presence, models reach F1 ≈ 0.98, which is near ceiling. However, the same experiments report that this performance reflects dataset bias, since most comics are humorous and models have nearly random performance distinguishing genuinely non-humorous ones (Ryan et al., 12 Sep 2025). By contrast, the more diagnostic tasks remain difficult.

On Sound Effect Detection, GPT-4o achieves F1: 0.82, while typical open models score 0.70–0.74. On Panel Contribution, the best reported result is GPT-4o with F1: 0.77, whereas open models are in the 0.49–0.54 range. On Modality Attribution, GPT-4o reaches F1: 0.63, while open models span 0.21–0.58 (Ryan et al., 12 Sep 2025).

The Humor Style task is harder still. The reported best total F1 is 0.50 for GPT-4o, while open models range from 0.09 (LLaVA) to 0.48 (Gemini) in the synopsis table (Ryan et al., 12 Sep 2025). The qualitative account states that models perform best on highly visual/personification jokes, while sarcasm, dark, and subtle humor types are poorly detected.

Narrative ordering remains a major bottleneck. On Sequence Recognition, the best listed model is Gemini-1.5-Pro with Panel Acc: 64.5%, whereas typical open LMMs achieve 31–42%, and the measured human baseline is 100% on the tested subset (Ryan et al., 12 Sep 2025). This gap is central to PixelHumor’s scientific value: it shows that high-capacity multimodal systems still struggle with the sequential structure that underpins comic timing.

For Humor Interpretation, GPT-4o obtains a mean of 5.8/7, while open models fall in the 3.0–4.6/7 range. Yet even here, human-written explanations are consistently preferred: in 69% of rating cases, humans selected the human explanation over any model output (Ryan et al., 12 Sep 2025). This indicates that current LMM explanations often sound plausible without capturing the specific mechanism of the joke.

5. Failure modes, diagnostic value, and common misconceptions

PixelHumor’s central empirical conclusion is that surface-level recognition is nearly solved, whereas true comprehension is largely unsolved (Ryan et al., 12 Sep 2025). The benchmark therefore separates two questions that are often conflated: whether a model can identify that a comic is intended as humor, and whether it can explain why the humor works.

Several recurrent failure modes are reported. One is heuristic bias: models default to “both modalities” for modality attribution or assume a canonical left-to-right reading order even when the panel order has been randomized. Another is hallucination, in which models generate plausible but factually incorrect explanations. A third is reliance on template explanations, such as generic statements about “unexpected twist” or “absurd situation,” without pinpointing the relevant narrative-visual interaction (Ryan et al., 12 Sep 2025).

Performance also degrades with comic length. The benchmark reports that longer comics produce weaker interpretation performance, which indicates difficulty with extended narrative/factual tracking across multiple panels (Ryan et al., 12 Sep 2025). This is consistent with the broader implication drawn in the paper: LMMs need better hierarchical/context tracking to maintain comedic buildup and punchline delivery.

PixelHumor also highlights a misconception about multimodal scaling. Although GPT-4o and Gemini consistently outperform open models on nontrivial tasks, the benchmark concludes that increasing general multimodal capability has not closed the gap on contextual integration, cultural and social nuance, narrative reasoning, and human alignment (Ryan et al., 12 Sep 2025). The benchmark’s significance lies precisely in making these deficits measurable.

6. Position within computational humor research

PixelHumor occupies a specific niche within a larger humor-research landscape. It differs from HumorDB, which is an image-only dataset of 3,545 images and 1,271 pairs designed for binary classification, range regression, and pairwise comparison on graphical humor; HumorDB emphasizes subtle visual cues and paired contrasts, whereas PixelHumor centers multi-panel narrative and multimodal comic reasoning (Jain et al., 2024). It also differs from OxfordTVG-HIC, which provides approximately 2.9M image-text pairs with humour scores to support humorous caption generation, including a position-conditioned loss for captioning models (Li et al., 2023), and from the earlier Neural Joking Machine, which used ResNet-152 + LSTM, Funny Score, and BoketeDB for humorous image captioning (Yoshida et al., 2018).

In video, adjacent benchmarks and datasets address complementary phenomena. ExFunTube contains 10,136 user-generated short-form funny videos with timestamps and detailed human-written explanations, emphasizing humor arising from the interaction of visual and verbal cues (Ko et al., 2023). CVLA targets short-form video humor detection through comment-aided video-language alignment and multi-modal contrastive pre-training on DY11k and UR-FUNNY (Liu et al., 2024). V-HUB is a visual-centric humor understanding benchmark for video LLMs built from 960 videos, mostly with visual-centric humor, and shows that all tested models suffer a clear degradation when moving from text-based to raw video evaluation (Shi et al., 30 Sep 2025).

PixelHumor is also related to recent work on internet memes and online communities. The study of r/ProgrammerHumor reports that image-based submissions are significantly more likely to achieve higher Reddit scores than text-only submissions, with 80% of annotated images adding context to the humor, which reinforces the broader relevance of multimodal humor analysis beyond comics (Kuutila et al., 2024). At the generation end, the HUMOR framework treats meme creation as a multimodal problem requiring hierarchical, multi-path Chain-of-Thought, group-wise pairwise reward modeling, and group-wise reinforcement learning aligned to within-template human preferences (Li et al., 31 Dec 2025).

Within this landscape, PixelHumor’s distinctive contribution is to make comic-specific multimodal humor understanding measurable at the intersection of narrative sequencing, style recognition, punchline localization, and explanatory reasoning. This suggests that PixelHumor is not merely another humor dataset, but a benchmark for probing whether LMMs can integrate visual and textual evidence across time-like panel structure in a way that approximates human comic comprehension (Ryan et al., 12 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to PixelHumor.