Papers
Topics
Authors
Recent
Search
2000 character limit reached

Citation Farming on ResearchGate: Blatant and Effective

Published 15 Apr 2026 in cs.SI | (2604.13784v2)

Abstract: We investigate platform-native citation farming on ResearchGate by analyzing almost 3000 papers uploaded by five suspected boosting-service provider accounts. From the uploaded papers and associated metadata, we construct both paper-level and author-level citation networks. We introduce an interpretable structural signal for coordinated boosting, equal references groups: clusters of papers with equal reference lists. We find that many papers from our collection exhibit this motif, that is, they disproportionately cite a small set of authors, consistent with coordinated or automated boosting rather than independent scholarly practice. Finally, we show that for some authors in our dataset a substantial share of their citations can be attributed to these suspicious groups. A different citation network was used to validate the rareness of such motifs in legitimate scientific work.

Summary

  • The paper demonstrates that coordinated citation farming on ResearchGate artificially inflates metrics, undermining the reliability of scholarly evaluation.
  • It employs an actor-seeded, motif-based network analysis to identify suspicious patterns such as equal references groups among over 3,000 papers.
  • Results show that for some authors, over 80% of citations come from manipulated groups, highlighting the urgent need for robust countermeasures.

Citation Farming on ResearchGate: Analytical Summary

Introduction and Context

The paper "Citation Farming on ResearchGate: Blatant and Effective" (2604.13784) presents an empirical, network-based study of coordinated citation inflation activities occurring natively within ResearchGate. The authors focus on a forensic, actor-seeded investigation workflow, starting from suspected boosting-service provider accounts (SSPAs) rather than anomaly scanning of entire corpora. They contextualize their study within the burgeoning proliferation of LLM-generated publications, facilitated by unregulated upload practices on platforms such as ResearchGate, and highlight the potential pollution and distortion of scholarly metrics as a result. This research builds on prior work, including both structural and semantic anomaly detection methods [Avros2023, Liu2022], and cites previous empirical validations of citation manipulation phenomena [ibrahim_citation_2025, KIRILOVA2025101604].

Data Collection and Methodology

Leveraging manual actor-seeded discovery, the authors identify five SSPAs publishing almost 3,000 papers—primarily preprints with unverifiable co-authors—on ResearchGate. Metadata extraction from ResearchGate profiles and publications enables construction of directed citation networks at both the author and paper level. The dataset comprises nearly 13,000 cited articles and over 22,000 authors—a subset of whom are uniquely identifiable via persistent platform links. This approach circumvents unreliable PDF parsing and exploits the consistency of ResearchGate's metadata structure. The resulting dataset, a valuable resource for subsequent analysis, is made publicly available. Figure 1

Figure 1: Timeline of illegitimate account publications showing sustained activity and author count surge post-2022.

Analysis reveals two paper archetypes: pre-2022 papers with numerous authors—often legitimate, yet re-uploaded for co-authorship claims—and post-2022 synthetic papers, usually with minimal or non-existent author information, highly repetitive filenames, sparse and biased reference lists, and often incoherent textual structure. The shift to higher-frequency, more blatant bibliometric manipulation aligns temporally with the widespread availability of LLM-generated content.

Structural Signals: Equal References Groups

Central to the detection framework is the concept of equal references groups—maximal cliques of papers with identically structured reference lists, frequently citing the same set of authors. These motifs represent interpretable, graph-based signals indicative of coordinated citation inflation. Figure 2

Figure 2

Figure 2: Visualization of equal references groups, illustrating green citing papers sharing identical outgoing citations to red cited papers.

Such motifs are anomalous when compared to genuine scholarly practice, wherein even closely related works exhibit substantial citation list differences. The presence of recurring citation structures oriented toward a core beneficiary set is a strong indicator of manipulation. Through motif enumeration, the authors identify 240 distinct groups involving multiple papers with identical reference lists.

To assess the beneficiaries, the study examines the frequency and coverage of cited authors across these structural motifs. Authors repeatedly cited in numerous groups, especially with multiple distinct papers, emerge as principal recipients of artificial bibliometric inflation. Figure 3

Figure 3: Scatter plot quantifying the correlation between distinct cited papers per author and the number of motifs they appear in, revealing prominent outliers as major citation beneficiaries.

Empirical results demonstrate that some authors receive over 80% of their ResearchGate citations from such suspicious motifs—a quantitative finding which underscores the disproportionate impact of citation farming.

Comparative Analysis

For validation, motif analysis is replicated on the HepPh citation network [gehrke2003overview], comprising 35,000 legitimate papers in high energy physics phenomenology. Results show that while HepPh contains more motifs in raw count, the sizes and overlap of these groups are substantively less pronounced compared to those in ResearchGate's manipulated network. After normalization, the frequency and density of equal references groups in the ResearchGate sample is an outlier, with larger groups and more frequent repetitions linked to boosting-service activity. Figure 4

Figure 4: Comparative plot illustrating motif group sizes and counts in HepPh versus ResearchGate, demonstrating the abnormal prevalence and structure in the latter.

Discussion and Implications

This investigation confirms the existence and efficacy of platform-native citation boosting, facilitated by unregulated preprint uploads and minimal user verification on ResearchGate. Motifs such as equal references groups are powerful structural signals for identifying both perpetrators and beneficiaries. The study quantifies bibliometric distortion at an unprecedented level, evidencing that for some authors, a substantial fraction of citation metrics is entirely attributable to suspicious activity.

The practical implications are profound: citation farming directly undermines the validity of bibliometric indicators traditionally used in faculty evaluation and academic reputation assessment. Theoretical implications center on the fragility of metric-based recognition and the susceptibility of scholarly networks to coordinated manipulation—a challenge exacerbated by the proliferation of LLM-generated content and the opacity of preprint metadata policies.

The study suggests several avenues for future AI-driven work, including scalable motif detection in larger and cross-platform datasets, integration of semantic citation context models for enhanced anomaly detection [Liu2022], and the development of graph-based countermeasures for real-time monitoring of citation manipulation. The dataset released enables further exploration and refinement of these approaches, supporting broader studies in metric fraud and scientific integrity.

Conclusion

This paper provides conclusive evidence of coordinated citation farming on ResearchGate, leveraging structural citation network analysis to isolate and characterize manipulated bibliometric motifs. Quantitative findings highlight the substantial impact on both individual author metrics and broader network structure, validating concerns that current academic evaluation metrics are vulnerable to fraud. Comparative analysis with legitimate citation networks further strengthens the claim of abnormal motif prevalence. This work lays a foundation for both practical and theoretical advances in bibliometric anomaly detection and underscores the necessity for robust platform governance in the age of LLM-generated scientific content.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.