Zero-Shot Activity Recognition with Videos (2002.02265v1)

Published 22 Jan 2020 in cs.CV, cs.CL, cs.LG, and stat.ML

Abstract: In this paper, we examined the zero-shot activity recognition task with the usage of videos. We introduce an auto-encoder based model to construct a multimodal joint embedding space between the visual and textual manifolds. On the visual side, we used activity videos and a state-of-the-art 3D convolutional action recognition network to extract the features. On the textual side, we worked with GloVe word embeddings. The zero-shot recognition results are evaluated by top-n accuracy. Then, the manifold learning ability is measured by mean Nearest Neighbor Overlap. In the end, we provide an extensive discussion over the results and the future directions.

Citations (1)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Follow-up Questions

We haven't generated follow-up questions for this paper yet.

Generate Now

Authors (1)

Evin Pinar Ornek

Zero-Shot Activity Recognition with Videos (2002.02265v1)

Summary

Follow-up Questions

Related Papers

Authors (1)