Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Figure Captioning with Reasoning and Sequence-Level Training (1906.02850v1)

Published 7 Jun 2019 in cs.CV and cs.CL

Abstract: Figures, such as bar charts, pie charts, and line plots, are widely used to convey important information in a concise format. They are usually human-friendly but difficult for computers to process automatically. In this work, we investigate the problem of figure captioning where the goal is to automatically generate a natural language description of the figure. While natural image captioning has been studied extensively, figure captioning has received relatively little attention and remains a challenging problem. First, we introduce a new dataset for figure captioning, FigCAP, based on FigureQA. Second, we propose two novel attention mechanisms. To achieve accurate generation of labels in figures, we propose Label Maps Attention. To model the relations between figure labels, we propose Relation Maps Attention. Third, we use sequence-level training with reinforcement learning in order to directly optimizes evaluation metrics, which alleviates the exposure bias issue and further improves the models in generating long captions. Extensive experiments show that the proposed method outperforms the baselines, thus demonstrating a significant potential for the automatic captioning of vast repositories of figures.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Charles Chen (8 papers)
  2. Ruiyi Zhang (98 papers)
  3. Eunyee Koh (36 papers)
  4. Sungchul Kim (65 papers)
  5. Scott Cohen (40 papers)
  6. Tong Yu (119 papers)
  7. Ryan Rossi (67 papers)
  8. Razvan Bunescu (17 papers)
Citations (36)

Summary

We haven't generated a summary for this paper yet.