Papers
Topics
Authors
Recent
Search
2000 character limit reached

PST-Bench: Tracing and Benchmarking the Source of Publications

Published 25 Feb 2024 in cs.DL and cs.CL | (2402.16009v1)

Abstract: Tracing the source of research papers is a fundamental yet challenging task for researchers. The billion-scale citation relations between papers hinder researchers from understanding the evolution of science efficiently. To date, there is still a lack of an accurate and scalable dataset constructed by professional researchers to identify the direct source of their studied papers, based on which automatic algorithms can be developed to expand the evolutionary knowledge of science. In this paper, we study the problem of paper source tracing (PST) and construct a high-quality and ever-increasing dataset PST-Bench in computer science. Based on PST-Bench, we reveal several intriguing discoveries, such as the differing evolution patterns across various topics. An exploration of various methods underscores the hardness of PST-Bench, pinpointing potential directions on this topic. The dataset and codes have been available at https://github.com/THUDM/paper-source-trace.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (39)
  1. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. Anthropic. 2023. Introducing claude. https://www.anthropic.com/news/introducing-claude.
  3. Scibert: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pages 3615–3620.
  4. Leo Breiman. 2001. Random forests. Machine learning, 45:5–32.
  5. Cogdl: A comprehensive library for graph deep learning. In Proceedings of the ACM Web Conference 2023, pages 747–758.
  6. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255.
  7. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4171–4186.
  8. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  9. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  10. To cite, or not to cite? detecting citation contexts in text. In Advances in Information Retrieval: 40th European Conference on IR Research, pages 598–603.
  11. Science of science. Science, 359(6379):eaao0185.
  12. Identifying important citations using contextual information from full text. In 2017 ACM/IEEE Joint Conference on Digital Libraries, pages 1–8.
  13. Detecting topic evolution in scientific literature: how can citations help? In Proceedings of the 18th ACM conference on Information and knowledge management, pages 957–966.
  14. Xiaorui Jiang and Jingqiang Chen. 2023. Contextualised segment-wise citation function classification. Scientometrics, 128(9):5117–5158.
  15. Measuring the evolution of a scientific field through citation frames. Transactions of the Association for Computational Linguistics, 6:391–406.
  16. Thomas N Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations.
  17. Ideareader: A machine reading system for understanding the idea flow of scientific publications. arXiv preprint arXiv:2209.13243.
  18. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022.
  19. Saurav Manchanda and George Karypis. 2021. Evaluating scholarly impact: Towards content-aware bibliometrics. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6041–6053.
  20. OpenAI. 2022. Introducing chatgpt. https://openai.com/blog/chatgpt.
  21. Scikit-learn: Machine learning in python. the Journal of machine Learning research, 12:2825–2830.
  22. Andrea Pellegrini. 2021. Arm neoverse n2: Arm’s 2 nd generation high performance infrastructure cpus and system ips. In 2021 IEEE Hot Chips 33 Symposium, pages 1–27.
  23. David Pride and Petr Knoth. 2017. Incidental or influential?-challenges in automatically detecting citation importance using publication full texts. In Research and Advanced Technology for Digital Libraries: 21st International Conference on Theory and Practice of Digital Libraries, pages 572–578.
  24. Netsmf: Large-scale network embedding as sparse matrix factorization. In The World Wide Web Conference, pages 1509–1520.
  25. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763.
  26. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  27. André Seznec and Pierre Michaud. 2006. A case for (partially) tagged geometric history length branch prediction. The Journal of Instruction-Level Parallelism, 8:23.
  28. Tracing the evolution of ai in the past decade and forecasting the emerging trends. Expert Systems with Applications, 209:118221.
  29. The amd “zen 2” processor. IEEE Micro, 40(2):45–52.
  30. Line: Large-scale information network embedding. In Proceedings of the 24th international conference on world wide web, pages 1067–1077.
  31. Topic distributions over links on web. In 9th IEEE International Conference on Data Mining, pages 1010–1015.
  32. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085.
  33. Identifying meaningful citations. In AAAI workshop: Scholarly big data, volume 15, page 13.
  34. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  35. Graph Attention Networks. In Proceedings of the 6th International Conference on Learning Representations.
  36. Mrt: Tracing the evolution of scientific publications. IEEE Transactions on Knowledge and Data Engineering, 35(1):711–724.
  37. Mining algorithm roadmap in scientific publications. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1083–1092.
  38. OAG: Toward linking large-scale heterogeneous entity graphs. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 2585–2595.
  39. Prone: fast and scalable network representation learning. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, pages 4278–4284.
Citations (5)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.