EROS: Entity-Driven Controlled Policy Document Summarization
Abstract: Privacy policy documents have a crucial role in educating individuals about the collection, usage, and protection of users' personal data by organizations. However, they are notorious for their lengthy, complex, and convoluted language especially involving privacy-related entities. Hence, they pose a significant challenge to users who attempt to comprehend organization's data usage policy. In this paper, we propose to enhance the interpretability and readability of policy documents by using controlled abstractive summarization -- we enforce the generated summaries to include critical privacy-related entities (e.g., data and medium) and organization's rationale (e.g.,target and reason) in collecting those entities. To achieve this, we develop PD-Sum, a policy-document summarization dataset with marked privacy-related entity labels. Our proposed model, EROS, identifies critical entities through a span-based entity extraction model and employs them to control the information content of the summaries using proximal policy optimization (PPO). Comparison shows encouraging improvement over various baselines. Furthermore, we furnish qualitative and human evaluations to establish the efficacy of EROS.
- Privacy policies over time: Curation and analysis of a million-document dataset. In Proceedings of the Web Conference 2021, pages 2165–2176.
- Natural language processing with Python: analyzing text with the natural language toolkit. " O’Reilly Media, Inc.".
- Automated extraction and presentation of data practices in privacy policies. Proc. Priv. Enhancing Technol., 2021(2):88–110.
- bert2BERT: Towards reusable pretrained language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2134–2148, Dublin, Ireland. Association for Computational Linguistics.
- Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
- Hierarchical transformers for long document summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics.
- Markus Eberts and Adrian Ulges. 2020. Span-based joint entity and relation extraction with transformer pre-training.
- SpanNER: Named entity re-/recognition as span prediction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7183–7195, Online. Association for Computational Linguistics.
- Using question answering rewards to improve abstractive summarization. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 518–526.
- Ctrlsum: Towards generic controllable text summarization. ArXiv, abs/2012.04281.
- Teaching machines to read and comprehend. Advances in neural information processing systems, 28.
- Enumeration of extractive oracle summaries. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pages 386–396, Valencia, Spain. Association for Computational Linguistics.
- Spanbert: Improving pre-training by representing and predicting spans. Transactions of the Association for Computational Linguistics, 8:64–77.
- Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.
- A unified MRC framework for named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5849–5859, Online. Association for Computational Linguistics.
- Triggerner: Learning with entity triggers as explanations for named entity recognition. CoRR, abs/2004.07493.
- Learning to summarize from human feedback. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.
- Presumm: A bert-based unsupervised text summarization model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.
- Zhengyuan Liu and Nancy F. Chen. 2021. Controllable neural dialogue summarization with personal named entity planning. In Conference on Empirical Methods in Natural Language Processing.
- Exploring the limits of transfer learning with a unified text-to-text transformer.
- Reinforcement learning for bandit neural machine translation with simulated human feedback. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing.
- Abstractive summarization with combination of pre-trained sequence-to-sequence and saliency models. arXiv preprint arXiv:2003.13028.
- Proximal policy optimization algorithms. CoRR, abs/1707.06347.
- Bigpatent: A large-scale dataset for abstractive and coherent summarization. arXiv preprint arXiv:1906.03741.
- Privacy at scale: Introducing the PrivaSeer corpus of web privacy policies. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6829–6839, Online. Association for Computational Linguistics.
- Improving abstractive document summarization with salient information modeling. In Proceedings of the 27th International Conference on Computational Linguistics.
- A hierarchical reinforced sequence operation method for abstractive summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing.
- Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2103.05447.
- Named entity recognition as dependency parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6470–6476, Online. Association for Computational Linguistics.
- Pegasus: Pre-training with extracted gap-sentences for abstractive summarization.
- Macsum: Controllable summarization with mixed attributes.
- Zexuan Zhong and Danqi Chen. 2020. A frustratingly easy approach for entity and relation extraction. arXiv preprint arXiv:2010.12812.
- Enwei Zhu and Jinpeng Li. 2022. Boundary smoothing for named entity recognition. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7096–7108, Dublin, Ireland. Association for Computational Linguistics.
Paper Prompts
Sign up for free to create and run prompts on this paper.