CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
Abstract: This paper explores the challenges posed by aspect-based sentiment classification (ABSC) within pretrained LLMs (PLMs), with a particular focus on contextualization and hallucination issues. In order to tackle these challenges, we introduce CARBD-Ko (a Contextually Annotated Review Benchmark Dataset for Aspect-Based Sentiment Classification in Korean), a benchmark dataset that incorporates aspects and dual-tagged polarities to distinguish between aspect-specific and aspect-agnostic sentiment classification. The dataset consists of sentences annotated with specific aspects, aspect polarity, aspect-agnostic polarity, and the intensity of aspects. To address the issue of dual-tagged aspect polarities, we propose a novel approach employing a Siamese Network. Our experimental findings highlight the inherent difficulties in accurately predicting dual-polarities and underscore the significance of contextualized sentiment analysis models. The CARBD-Ko dataset serves as a valuable resource for future research endeavors in aspect-level sentiment classification.
- Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. URL https://arxiv.org/pdf/1810.04805.pdf.
- Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019. URL https://arxiv.org/pdf/1906.08237.
- BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online, July 2020. Association for Computational Linguistics. doi:10.18653/v1/2020.acl-main.703. URL https://aclanthology.org/2020.acl-main.703.
- Utilizing bert for aspect-based sentiment analysis via constructing auxiliary sentence. arXiv preprint arXiv:1903.09588, 2019. URL https://aclanthology.org/N19-1035.pdf.
- UniMSE: Towards unified multimodal sentiment analysis and emotion recognition. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7837–7851, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.534.
- A unified generative framework for aspect-based sentiment analysis. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2416–2429, Online, August 2021. Association for Computational Linguistics. doi:10.18653/v1/2021.acl-long.188. URL https://aclanthology.org/2021.acl-long.188.
- Learning implicit sentiment in aspect-based sentiment analysis with supervised contrastive pre-training. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 246–256, Online and Punta Cana, Dominican Republic, November 2021a. Association for Computational Linguistics. doi:10.18653/v1/2021.emnlp-main.22. URL https://aclanthology.org/2021.emnlp-main.22.
- Bing Liu. Sentiment analysis and opinion mining. Synthesis lectures on human language technologies, 5(1):1–167, 2012. URL https://www.cs.uic.edu/~liub/FBS/SentimentAnalysis-and-OpinionMining.pdf.
- Beta distribution guided aspect-aware graph for aspect category sentiment analysis with affective knowledge. In Proceedings of the 2021 conference on empirical methods in natural language processing, pages 208–218, 2021. URL https://aclanthology.org/2021.emnlp-main.19.pdf.
- Learning implicit sentiment in aspect-based sentiment analysis with supervised contrastive pre-training. arXiv preprint arXiv:2111.02194, 2021b. URL https://aclanthology.org/2021.emnlp-main.22.pdf.
- Indonlg: Benchmark and resources for evaluating indonesian natural language generation. arXiv preprint arXiv:2104.08200, 2021. URL https://aclanthology.org/2021.emnlp-main.699.pdf.
- The flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538, 2022. URL https://arxiv.org/pdf/2106.03193.pdf.
- Sentence-bert: Sentence embeddings using siamese bert-networks, 2019. URL https://arxiv.org/pdf/1908.10084.pdf.
- Emotions are universal: Learning sentiment based representations of resource-poor languages using siamese networks. arXiv preprint arXiv:1804.00805, 2018a. URL https://arxiv.org/pdf/1804.00805.pdf.
- Sentiment analysis of code-mixed languages leveraging resource rich languages. arXiv preprint arXiv:1804.00806, 2018b. URL https://arxiv.org/pdf/1804.00806.pdf.
- Siamese network-based supervised topic modeling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4652–4662, 2018. URL https://aclanthology.org/D18-1494.pdf.
- Kasn: Knowledge-aware siamese network for sentiment analysis. In AIIPCC 2022; The Third International Conference on Artificial Intelligence, Information Processing and Cloud Computing, pages 1–8. VDE, 2022. URL https://ieeexplore.ieee.org/document/10025876.
- Jangwon Park. Koelectra: Pretrained electra model for korean. https://github.com/monologg/KoELECTRA, 2020.
- Kr-electra: a korean-based electra model. https://github.com/snunlp/KR-ELECTRA, 2022.
- Junbum Lee. Kcelectra: Korean comments electra. https://github.com/Beomi/KcELECTRA, 2021.
- Kr-bert: A small-scale korean-specific language model. ArXiv, abs/2008.03979, 2020. URL https://arxiv.org/pdf/2008.03979.pdf.
- Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116, 2019. URL https://aclanthology.org/2020.acl-main.747.pdf.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.