Papers
Topics
Authors
Recent
Search
2000 character limit reached

HAUSER: Towards Holistic and Automatic Evaluation of Simile Generation

Published 13 Jun 2023 in cs.CL | (2306.07554v1)

Abstract: Similes play an imperative role in creative writing such as story and dialogue generation. Proper evaluation metrics are like a beacon guiding the research of simile generation (SG). However, it remains under-explored as to what criteria should be considered, how to quantify each criterion into metrics, and whether the metrics are effective for comprehensive, efficient, and reliable SG evaluation. To address the issues, we establish HAUSER, a holistic and automatic evaluation system for the SG task, which consists of five criteria from three perspectives and automatic metrics for each criterion. Through extensive experiments, we verify that our metrics are significantly more correlated with human ratings from each perspective compared with prior automatic metrics.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (38)
  1. Catherine Addison. 2001. “so stretched out huge in length”: Reading the extended simile. Style, 35(3):498–516.
  2. Comet: Commonsense transformers for automatic knowledge graph construction. arXiv preprint arXiv:1906.05317.
  3. Evaluation of text generation: A survey. arXiv preprint arXiv:2006.14799.
  4. It’s not rocket science: Interpreting figurative language in narratives. Transactions of the Association for Computational Linguistics, 10:589–606.
  5. Generating similes effortlessly like a pro: A style transfer approach for simile generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6455–6469.
  6. Probing simile knowledge from pre-trained language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5875–5887.
  7. Of human criteria and automatic metrics: A benchmark of the evaluation of story generation. In 29th International Conference on Computational Linguistics (COLING 2022).
  8. Michael Denkowski and Alon Lavie. 2014. Meteor universal: Language specific translation evaluation for any target language. In Proceedings of the ninth workshop on statistical machine translation, pages 376–380.
  9. Handling divergent reference texts when evaluating table-to-text generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4884–4895.
  10. Patrick Hanks. 2013. Lexical analysis: Norms and exploitations. Mit Press.
  11. Maps-kb: A million-scale probabilistic simile knowledge base. arXiv preprint arXiv:2212.05254.
  12. Lara L Jones and Zachary Estes. 2006. Roosters, robins, and alarm clocks: Aptness and conventionality in metaphor comprehension. Journal of Memory and Language, 55(1):18–32.
  13. Huiyuan Lai and Malvina Nissim. 2022. Multi-figurative language generation. In Proceedings of the 29th International Conference on Computational Linguistics, pages 5939–5954.
  14. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  15. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110–119.
  16. Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  17. Vlad Niculae and Cristian Danescu-Niculescu-Mizil. 2014. Brighter than gold: Figurative language in user generated comparisons. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 2008–2018.
  18. Towards holistic and automatic evaluation of open-domain dialogue generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3619–3629.
  19. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  20. Anthony M Paul. 1970. Figurative language. Philosophy & Rhetoric, pages 225–248.
  21. Russell S Pierce and Dan L Chiappe. 2008. The roles of aptness, conventionality, and working memory in the production of metaphors and similes. Metaphor and symbol, 24(1):1–19.
  22. Learning to recognize affective polarity in similes. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 190–200.
  23. Codebleu: a method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297.
  24. Carlos Roncero and Roberto G de Almeida. 2015. Semantic properties, aptness, familiarity, conventionality, and interpretive diversity scores for 84 metaphors and similes. Behavior research methods, 47(3):800–812.
  25. A survey of evaluation metrics used for nlg systems. ACM Computing Surveys (CSUR), 55(2):1–39.
  26. Metaphoric paraphrase generation. arXiv preprint arXiv:2002.12854.
  27. Ruber: An unsupervised method for automatic evaluation of open-domain dialog systems. In Thirty-Second AAAI Conference on Artificial Intelligence.
  28. Roi Tartakovsky and Yeshayahu Shen. 2018. ‘simple as a fire’: Making sense of the non-standard poetic simile. Journal of Literary Semantics, 47(2):103–119.
  29. Guy Tevet and Jonathan Berant. 2021. Evaluating the evaluation of diversity in natural language generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 326–346.
  30. Amos Tversky. 1977. Features of similarity. Psychological review, 84(4):327.
  31. Glue: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355.
  32. Meta4meaning: Automatic metaphor interpretation using corpus-derived word associations. In Proceedings of the Seventh International Conference on Computational Creativity. Sony CSL Paris.
  33. Writing polishment with simile: Task, dataset and a neural approach. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 14383–14392.
  34. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  35. Yunxiang Zhang and Xiaojun Wan. 2021. Mover: Mask, over-generate and rank for hyperbole generation. arXiv preprint arXiv:2109.07726.
  36. Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 563–578.
  37. “love is as complex as math”: Metaphor generation system for social chatbot. In Workshop on Chinese Lexical Semantics, pages 337–347. Springer.
  38. Texygen: A benchmarking platform for text generation models. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, pages 1097–1100.
Citations (4)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.