Multimodal Explanations by Predicting Counterfactuality in Videos (1812.01263v2)

Published 4 Dec 2018 in cs.CV

Abstract: This study addresses generating counterfactual explanations with multimodal information. Our goal is not only to classify a video into a specific category, but also to provide explanations on why it is not categorized to a specific class with combinations of visual-linguistic information. Requirements that the expected output should satisfy are referred to as counterfactuality in this paper: (1) Compatibility of visual-linguistic explanations, and (2) Positiveness/negativeness for the specific positive/negative class. Exploiting a spatio-temporal region (tube) and an attribute as visual and linguistic explanations respectively, the explanation model is trained to predict the counterfactuality for possible combinations of multimodal information in a post-hoc manner. The optimization problem, which appears during training/inference, can be efficiently solved by inserting a novel neural network layer, namely the maximum subpath layer. We demonstrated the effectiveness of this method by comparison with a baseline of the action recognition datasets extended for this task. Moreover, we provide information-theoretical insight into the proposed method.

Authors (4)

Atsushi Kanehira (24 papers)
Kentaro Takemoto (5 papers)
Sho Inayoshi (2 papers)
Tatsuya Harada (142 papers)

Citations (34)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Multimodal Explanations by Predicting Counterfactuality in Videos (1812.01263v2)

Summary

Related Papers