Papers
Topics
Authors
Recent
Search
2000 character limit reached

ComicScene154: Annotated Comics Scene Dataset

Updated 9 July 2026
  • ComicScene154 is a manually annotated dataset for scene segmentation in comics, defining scene-level narrative arcs across panels.
  • It leverages public-domain Golden Age Western comics to analyze narrative structure by marking scene boundaries in sequential panels.
  • A two-step LLM-based pipeline demonstrates that multimodal reasoning can roughly segment scenes, though results remain far from human agreement.

Searching arXiv for the target paper and closely related comics benchmarks to ground the article in current literature. ComicScene154 is a manually annotated dataset for scene segmentation in comics, introduced to support computational narrative analysis in a medium that is inherently multimodal and structurally discontinuous. Its central unit is the scene-level narrative arc: a higher-level semantic segment that is not reducible to page boundaries, isolated panels, or low-level layout structure. The dataset was designed from public-domain Western comics and paired with an initial scene segmentation benchmark based on multimodal prompting and refinement, with the explicit aim of extending comic analysis toward narrative structure rather than object- or text-level recognition alone (Paval et al., 22 Aug 2025).

1. Research problem and conceptual framing

ComicScene154 was created to address a gap in comic analysis: many downstream tasks require narrative context, but commonly used processing units are poorly aligned with story structure. Whole-comic processing includes too much irrelevant material; page-level segmentation follows physical layout rather than narrative organization; and random samples of panels or pages are unsuitable for identifying scene boundaries because scenes depend on broader narrative development (Paval et al., 22 Aug 2025).

The dataset treats comics as a useful abstraction for multimodal narrative data. Comics combine images and text, but unlike video they are composed of discrete panels whose continuity is inferred by readers through closure. Within this framework, a scene is treated as a semantic narrative unit rather than a layout primitive. The stated motivation is not only comic-specific analysis, but also broader multimodal narrative research, since panel sequences can be interpreted as compact, sampled visual-narrative streams.

This framing places ComicScene154 close to work on semantic segmentation in discourse. The task is not to identify local syntactic or perceptual structure, but to recover semantically coherent units organized by plot development, cast continuity, and temporal or spatial coherence. A plausible implication is that scene segmentation in comics functions as a multimodal analogue of semantic text segmentation, but with stronger dependence on visual continuity and cross-panel inference.

2. Corpus design and source material

All source material comes from Comic Book Plus, a repository of public-domain comics. The corpus contains 4 public-domain comic magazines, 34 distinct stories, and 154 pages total. The source material is explicitly Western rather than manga, because the dataset excludes manga on the grounds that manga differs substantially from classic Western comics in storytelling conventions, structural layout, and linguistic features (Paval et al., 22 Aug 2025).

The corpus spans multiple genres and years, with a concentration in the Golden Age of comics. The years are summarized in the abstract and table as approximately 1942–1962, while the body text describes the period as “roughly 1940–1960.” The dataset therefore reflects older public-domain Western comics rather than contemporary comic production.

Source magazine Genre
Alley Oop, Vol. 1 Humor
Champ Comics, Vol. 24 Heroes
Treasure Comics, Vol. 6 Fantasy
Western Love, Vol. 4 Love

Genre diversity is part of the design rationale. The dataset aims to vary storytelling style and narrative form, not merely page layout. At the same time, this composition imposes a strong historical and cultural prior: Golden Age Western comics dominate, and the dataset does not attempt cross-tradition coverage.

3. Annotation model and representation of scenes

The annotation scheme is panel-centric. Before annotation, the authors extracted all panels from comic pages, stored their coordinates, and numbered them in reading order. This preprocessing is foundational, because scene segmentation depends on a consistent narrative sequence; missing or misordered panels would distort the segmentation target (Paval et al., 22 Aug 2025).

Each panel is assigned a Boolean label indicating whether it is the start of a new scene. Scenes are therefore represented through boundary markers rather than directly enumerated span identifiers. Operationally, a scene is recoverable as the interval beginning at a marked panel and ending immediately before the next marked boundary.

The dataset adopts a scene definition from Rao et al. (2020) and relates scenes to narrative arcs in Cohn (2013). In the paper’s formulation, a scene is a plot-based semantic unit in which an overarching task is pursued by a certain cast of characters and, in most cases, maintains temporal and spatial coherence. The paper also states that it treats a scene analogously to a narrative arc. In ComicScene154, “scene” and “scene-level narrative arc” are thus near-equivalent analytical units.

Annotation instructions were intentionally lightweight. Annotators were told to review comic images, identify where they perceived transitions between scenes, and mark each panel that signals the start of a new narrative arc. They received a brief introduction, the working scene definition, and one sample annotated example. The appendix emphasizes that segmentation is subjective and that deviations are acceptable so long as agreement persists on the “core panels” of a scene. This formalizes ambiguity rather than eliminating it.

4. Reliability, metric, and benchmark protocol

ComicScene154 includes an explicit reliability study. One-third of the dataset was independently labeled for evaluation by three different groups of two annotators, and those labels were compared against the authors’ annotations. The annotators were friends or colleagues of the authors, followed the written guidelines, received no additional input, and were unpaid volunteers (Paval et al., 22 Aug 2025).

Agreement is measured with pkp_k, adopted from semantic text segmentation literature. The paper describes pkp_k qualitatively as the proportion of sliding windows segmented inconsistently between two annotations, with pk=0p_k = 0 indicating perfect alignment and pk=1p_k = 1 total misalignment. The window size is fixed at k=3k = 3, chosen by averaging scene lengths from the authors’ annotations and the external annotators’ annotations.

Excerpt set summary Value
Tester 1 average pkp_k 0.15
Tester 2 average pkp_k 0.19
In-between average pkp_k 0.21
Average tester pkp_k 0.17

These numbers indicate meaningful but imperfect agreement. The paper interprets this as evidence that annotators share a broadly recognizable concept of scenes, while also revealing substantial subjectivity in boundary placement and scene granularity. One cited example is a Western Love excerpt in which one annotator preferred much shorter scenes than another. The main challenge is therefore not only modeling difficulty, but also the absence of a fully intersubjective definition of scene boundaries.

The benchmark does not define a conventional train/validation/test split. There is no supervised training protocol in the usual sense. Instead, the paper evaluates repeated prompted inference against the human annotations using pkp_k, reflecting the exploratory status of the baseline.

5. Baseline scene segmentation pipeline and empirical results

The baseline is a two-step LLM-based pipeline rather than a dedicated trained segmentation model. In the first step, a multimodal reasoning model receives the comic page, panel coordinates, and reading order, and is prompted to describe narrative arcs and identify the panel where each arc begins. In the second step, those initial outputs are refined by a reasoning-based LLM. The paper states that the same Gemini reasoning model was used for both stages: gemini-2.0-flash-thinking-exp (Paval et al., 22 Aug 2025).

The effective representation is minimal and prompt-based. Inputs consist of the comic page image, panel coordinates, panel reading order, and the scene-boundary prompt. No OCR pipeline, no explicit panel embedding model, and no handcrafted feature engineering are reported. The benchmark therefore measures prompted multimodal reasoning under structured page context, not learned scene segmentation.

The evaluation protocol addresses output stochasticity by repeated inference. For each comic, the initial stage is run 10 times. Each initial output is then refined 10 times, yielding 100 evaluations per comic for the second stage. The paper reports average performance, best refined performance associated with a single initial output, and “In-between” consistency across model runs.

Setting Average pkp_k0 In-between
Random segmentation 0.46 —
Initial multimodal prompting 0.42 0.06
Refined prompting, all iterations 0.39 0.05
Refined prompting, best iteration 0.34 —

The initial multimodal stage is only slightly better than random: 0.42 versus 0.46. Refinement improves the average to 0.39, with 0.34 in the best-performing refined setting. Even the best refined result remains substantially worse than human agreement levels. At the same time, the low “In-between” values indicate strong internal consistency across model outputs. The paper interprets this as evidence that the model is applying stable heuristics that do not align well with the human notion of a scene as a semantic narrative arc.

This result is central to the dataset’s significance. ComicScene154 is not presented as a solved benchmark seed; it is presented as a setting in which current multimodal systems remain near-random in the initial setup and only modestly improve through reasoning-based refinement.

6. Position within comics research, limitations, and implications

ComicScene154 occupies a different level of abstraction from the better-established comic benchmarks focused on panel-local or object-local tasks. CoMix is designed for multi-task comic understanding and centers object detection, speaker identification, character re-identification, reading order, character naming, and dialog generation rather than scene segmentation (Vivoli et al., 2024). MaRU addresses dialogue retrieval and scene retrieval at the frame or panel level, not scene-level narrative arcs spanning multiple panels (Shen et al., 2023). COMICS Text+ targets OCR quality in Western comics and improves text detection, recognition, and cloze-style downstream tasks, but it does not annotate narrative scenes (Soykan et al., 2022). “Panel Transitions for Genre Analysis in Visual Narratives” foregrounds inter-panel transition types such as Action-to-Action, Aspect-to-Aspect, Subject-to-Subject, and Scene-to-Scene, which are structurally adjacent to scene segmentation but serve genre analysis rather than boundary annotation (Chen et al., 2023).

This suggests that ComicScene154 fills a specific gap: it targets higher-level multimodal semantics organized over panel sequences rather than over single panels, text regions, or object instances. In that sense it complements, rather than replaces, OCR, retrieval, transition modeling, and panel-understanding benchmarks.

The dataset’s limitations are explicit. It contains 4 magazines, 34 stories, and 154 pages, which makes it small by contemporary machine learning standards. Its source domain is heavily skewed toward Golden Age Western comics; manga and other comic traditions are excluded; and transfer to modern comics is uncertain because artistic style, narrative complexity, pacing, paneling conventions, and dialogue structure may differ. Annotation subjectivity is intrinsic to the task, and annotator background may introduce bias because the external annotators were friends or colleagues of the authors and were unpaid volunteers (Paval et al., 22 Aug 2025).

Even with those limitations, the dataset has clear implications for multimodal narrative analysis. The paper explicitly connects it to story summarization, character identification, entity tracking, digital humanities uses such as analyzing storytelling patterns and styles, and possible transfer to movies and videos through sampled-frame representations. A plausible implication is that ComicScene154 is most valuable not as a large-scale training corpus, but as a focused challenge set for testing whether a system can move beyond panel recognition toward scene-level narrative organization.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ComicScene154.