Temporal Validity Change Prediction
Abstract: Temporal validity is an important property of text that is useful for many downstream applications, such as recommender systems, conversational AI, or story understanding. Existing benchmarking tasks often require models to identify the temporal validity duration of a single statement. However, in many cases, additional contextual information, such as sentences in a story or posts on a social media profile, can be collected from the available text stream. This contextual information may greatly alter the duration for which a statement is expected to be valid. We propose Temporal Validity Change Prediction, a natural language processing task benchmarking the capability of machine learning models to detect contextual statements that induce such change. We create a dataset consisting of temporal target statements sourced from Twitter and crowdsource sample context statements. We then benchmark a set of transformer-based LLMs on our dataset. Finally, we experiment with temporal validity duration prediction as an auxiliary task to improve the performance of the state-of-the-art model.
- Predicting the occurrence of life events from user’s tweet history. In 2018 IEEE 12th International Conference on Semantic Computing (ICSC), pages 219–226. IEEE.
- Axel Almquist and Adam Jatowt. 2019. Towards content expiry date determination: Predicting validity periods of sentences. In European Conference on Information Retrieval, pages 86–101. Springer.
- Sitaram Asur and Bernardo A Huberman. 2010. Predicting the future with social media. In 2010 IEEE/WIC/ACM international conference on web intelligence and intelligent agent technology, volume 1, pages 492–499. IEEE.
- Prajjwal Bhargava and Vincent Ng. 2022. Commonsense knowledge reasoning and generation with pre-trained language models: a survey. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 12317–12325.
- Chatgpt is a knowledgeable but inexperienced solver: An investigation of commonsense problem in large language models. arXiv preprint arXiv:2303.16421.
- Signature verification using a" siamese" time delay neural network. Advances in neural information processing systems, 6.
- Mitigating reporting bias in semi-supervised temporal commonsense inference with probabilistic soft logic. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 10454–10462.
- Combining technical analysis with sentiment analysis for stock price prediction. In 2011 IEEE ninth international conference on dependable, autonomic and secure computing, pages 800–807. IEEE.
- Reasoning with transformer-based models: Deep learning, but shallow reasoning. In 3rd Conference on Automated Knowledge Base Construction.
- On the round number bias and wisdom of crowds in different response formats for numerical estimation. Scientific Reports, 12(1):1–18.
- Temporal natural language inference: Evidence-based evaluation of temporal text validity. Proceedings of the 45th European Conference on Information Retrieval (ECIR 2023), Springer LNCS.
- Marc W Howard. 2018. Memory as perception of the past: compressed time inmind and brain. Trends in cognitive sciences, 22(2):124–136.
- Adam Jatowt and Ching-man Au Yeung. 2011. Extracting collective expectations about the future from large text collections. In Proceedings of the 20th ACM international conference on Information and knowledge management, pages 1259–1264.
- Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, volume 1, page 2.
- Towards a language model for temporal commonsense reasoning. In Proceedings of the Student Research Workshop Associated with RANLP 2021, pages 78–84.
- Lifetime of tweets: a statistical analysis. Social Network Analysis and Mining, 12(1):101.
- Predicting iphone sales from iphone tweets. In 2014 IEEE 18th International Enterprise Distributed Object Computing Conference, pages 81–90. IEEE.
- Location inference for non-geotagged tweets in user timelines. IEEE Transactions on Knowledge and Data Engineering, 31(6):1150–1165.
- Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
- Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. In International Conference on Learning Representations.
- Integrating predictive analytics and social media. In 2014 IEEE Conference on Visual Analytics Science and Technology (VAST), pages 193–202. IEEE.
- Commonsense temporal action knowledge (cotak) dataset. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM 2023).
- Bill MacCartney. 2009. Natural language inference. Stanford University.
- James Manyika. 2023. An overview of bard: an early experiment with generative ai.
- A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 839–849.
- A survey on applications of siamese neural networks in computer vision. In 2020 International Conference for Emerging Technology (INCET), pages 1–5. IEEE.
- Torque: A reading comprehension dataset of temporal ordering questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1158–1172.
- Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
- Alice++: Adversarial training for robust and effective temporal reasoning. In Proceedings of the 35th Pacific Asia Conference on Language, Information and Computation, pages 373–382.
- Adversarial training for commonsense inference. In Proceedings of the 5th Workshop on Representation Learning for NLP (RepL4NLP-2020), pages 55–60. Association for Computational Linguistics.
- Timedial: Temporal commonsense reasoning in dialog. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7066–7076.
- Guy Rosin and Kira Radinsky. 2022. Temporal attention for language models. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1498–1508.
- Time masking for temporal language models. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, pages 833–841.
- Forecasting people’s action via social media data. In 2020 IEEE International Conference on Big Data (Big Data), pages 5254–5259. IEEE.
- Commonsense reasoning for natural language understanding: A survey of benchmarks, resources, and approaches. arXiv preprint arXiv:1904.01172, pages 1–60.
- Recent advances in natural language inference: A survey of benchmarks, resources, and approaches. arXiv preprint arXiv:1904.01172.
- Self-explaining structures improve nlp models. arXiv preprint arXiv:2012.01786.
- Hikaru Takemura and Keishi Tajima. 2012. Tweet classification based on their lifetime duration. In Proceedings of the 21st ACM international conference on Information and knowledge management, pages 2367–2370.
- Towards benchmarking and improving the temporal reasoning capability of large language models. arXiv preprint arXiv:2306.08952.
- Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
- Lav R Varshney and John Z Sun. 2013. Why do we perceive logarithmically? Significance, 10(1):28–31.
- Attention is all you need. Advances in neural information processing systems, 30.
- Bitimebert: Extending pre-trained language representations with bi-temporal information. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 812–821.
- Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
- Georg Wenzel and Adam Jatowt. 2023. An overview of temporal commonsense reasoning and acquisition. arXiv preprint arXiv:2308.00002.
- Improving event duration prediction via time-aware pre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3370–3378.
- Cocolm: Complex commonsense enhanced language model with discourse relations. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1175–1187.
- Reasoning about goals, steps, and temporal ordering with wikihow. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4630–4639.
- “going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics.
- Temporal common sense acquisition with minimal supervision. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7579–7589.
- Temporal reasoning on implicit events from distant supervision. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1361–1371.
- Generating temporally-ordered event sequences via event optimal transport. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1875–1884.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.