Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Automatically Categorising GitHub Repositories by Application Domain (2208.00269v1)

Published 30 Jul 2022 in cs.SE and cs.LG

Abstract: GitHub is the largest host of open source software on the Internet. This large, freely accessible database has attracted the attention of practitioners and researchers alike. But as GitHub's growth continues, it is becoming increasingly hard to navigate the plethora of repositories which span a wide range of domains. Past work has shown that taking the application domain into account is crucial for tasks such as predicting the popularity of a repository and reasoning about project quality. In this work, we build on a previously annotated dataset of 5,000 GitHub repositories to design an automated classifier for categorising repositories by their application domain. The classifier uses state-of-the-art natural language processing techniques and machine learning to learn from multiple data sources and catalogue repositories according to five application domains. We contribute with (1) an automated classifier that can assign popular repositories to each application domain with at least 70% precision, (2) an investigation of the approach's performance on less popular repositories, and (3) a practical application of this approach to answer how the adoption of software engineering practices differs across application domains. Our work aims to help the GitHub community identify repositories of interest and opens promising avenues for future work investigating differences between repositories from different application domains.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Francisco Zanartu (4 papers)
  2. Christoph Treude (137 papers)
  3. Bruno Cartaxo (9 papers)
  4. Hudson Silva Borges (2 papers)
  5. Pedro Moura (6 papers)
  6. Markus Wagner (90 papers)
  7. Gustavo Pinto (33 papers)
Citations (3)

Summary

We haven't generated a summary for this paper yet.