Paper Copilot: AI Peer Review Transparency
- The paper demonstrates that aggregating and visualizing AI/ML peer review data enhances transparency and accountability across multiple venues.
- The platform standardizes multi-source data using automated bots and community submissions, providing comprehensive analytical insights.
- The analysis reveals significant differences in reviewer confidence and engagement metrics, supporting calls for regulated, open peer review systems.
Paper Copilot is a research services platform built to aggregate and analyze AI/ML conference peer-review data and make open statistics available to the community. Its stated mission is to increase transparency, accountability, and community engagement around the review and decision process at top AI/ML venues, and to provide quantitative evidence to inform policy debates on peer-review models. Originating from concerns about rapidly growing submissions, often exceeding 10,000, and pressure on traditional review practices, the site was publicly launched two years before the 2025 position paper that used it to argue that the AI and machine learning community should adopt a more transparent, open, and well-regulated peer-review process (Yang, 2 Feb 2025).
1. Origins and stated objectives
Paper Copilot was introduced as a response to the scaling pressures of contemporary AI/ML conference reviewing. The platform’s goals are explicit: to collect and standardize multi-venue review data and timelines, visualize review score distributions, engagement dynamics, and author or affiliation patterns over time, quantify community interest in transparent reviewing through traffic and usage analytics, and support a position favoring more transparent, open, and well-regulated peer review (Yang, 2 Feb 2025).
In the position paper, the platform is not presented merely as a dashboard. It functions as an evidentiary instrument in a broader argument about governance of peer review. The paper links conference growth, reviewer load, and opacity in decision procedures to the need for shared statistics and public scrutiny. Paper Copilot is therefore framed simultaneously as infrastructure, a public-facing observatory, and a policy intervention.
The platform’s mission is closely tied to the distinction between review-platform adoption and review transparency. One of the paper’s central observations is that movement from traditional closed platforms such as Microsoft CMT to open platforms such as OpenReview does not by itself imply a more transparent disclosure regime. The paper’s figures on “Adoption of Review Platforms” and “Review Disclosure Preferences” are used to separate software migration from actual public visibility of reviews, discussions, and decision artifacts (Yang, 2 Feb 2025).
2. Data aggregation, platform design, and coverage
Paper Copilot aggregates data from official conference websites, review platforms, notably OpenReview, and community submissions. Its data aggregation pipeline cleans, merges, and stores data from multiple sources using standardized processing; resulting datasets are made open-source and visualized interactively. For venues that publish review artifacts, automated retrieval uses the OpenReview API to pull ratings, confidence levels, reviewer comments, and temporal profiles of score and discussion evolution, with bots running daily. Additional bots on official conference sites collect descriptive details such as author identities and affiliations and identify inconsistencies across sources (Yang, 2 Feb 2025).
For partially open or fully closed venues, the system supplements official data with voluntary community submissions through Google Forms. The position paper reports 3,876 valid community responses over the past year. This mechanism is especially important for venues such as ICML and CVPR, where official review artifacts remain private. The platform also tracks traffic and engagement metrics on its own statistics pages through Google Analytics, validated with Matomo, and states that no personally identifying information is collected (Yang, 2 Feb 2025).
The covered data include submissions and accepted paper metadata, review ratings and confidence, reviewer comments, author responses and discussion reply counts where public, decisions, timelines, author and affiliation metadata, and traffic and engagement metrics tied to statistics pages. The paper reports 10 years of available data processed from 24 venues across 9 subfields in AI/ML. Examples include ICLR, NeurIPS, CoRL, ICML, CVPR, ECCV, EMNLP, ACL, and KDD. The paper does not introduce formal mathematical definitions, formulas, or indices for transparency metrics; instead, it relies on descriptive statistics and visual analytics (Yang, 2 Feb 2025).
3. Taxonomy of peer-review openness
The position paper classifies AI/ML venues by disclosure mode while noting that all the compared venues share a double-blind review framework in which author and reviewer identities are hidden during reviewing. Openness is defined by the timing and extent of public disclosure of reviews, author responses and discussions, and decision notes (Yang, 2 Feb 2025).
| Model | Disclosure pattern | Examples |
|---|---|---|
| Fully open | Reviews, discussions, and responses are publicly visible in real time during the review phase | ICLR |
| Partially open | Reviews and discussions become public only after final decisions | NeurIPS, CoRL |
| Fully closed | Reviews, discussions, and decision notes remain private indefinitely | ICML, CVPR, ECCV |
This taxonomy is central to the paper’s argument because it distinguishes the public visibility of review artifacts from the mere use of an open-review platform. The figure on “Review Disclosure Preferences” indicates that true transparency levels remained mostly unchanged from 2015 to 2024 despite platform migrations, with 2025 excluded because of pending announcements (Yang, 2 Feb 2025).
The paper also associates each model with different oversight possibilities. Fully open systems permit real-time observation of score evolution and discussions; partially open systems preserve post-decision auditability while withholding review content during deliberation; fully closed systems maximize confidentiality but restrict external scrutiny. The article’s normative claim is not that one disclosure model is cost-free, but that openness requires explicit regulation rather than being treated as a purely technical platform choice.
4. Empirical findings from community usage and review dynamics
The platform is used in the paper as evidence of substantial international demand for review transparency. Over two years, Paper Copilot attracted over 200,000 active users, especially early-career researchers aged 18–34, from 177 countries, with a maximum daily peak of 15,000 unique visitors. The site recorded 6 million impressions, 1 million site views, and 4 million user-triggered events; in the preceding 28 days it received 50,000 organic clicks from Google Search. Organic search accounted for 59.9% of traffic, direct traffic 23.9%, referral traffic 9.4%, and organic social 7.6% (Yang, 2 Feb 2025).
The demographic and geographic distributions are also reported in detail. The 18–24 group is the largest age bracket; younger males have the longest average engagement time at 4 min 15 sec. User counts by country include 60,648 in the United States and 59,269 in China, while Singapore and Australia show engagement times above 3 minutes. Search rankings are especially strong for queries such as “ICLR 2025 statistics” and “NeurIPS 2024 accepted papers,” which the paper treats as evidence of organic demand for transparent review data (Yang, 2 Feb 2025).
Venue-level engagement further differentiates disclosure models. Click-through rates across venues fall between 66.08% and 86.49%, which the paper interprets as broad curiosity about review statistics regardless of venue transparency. ICLR, the fully open example, leads with 414,096 views, 88,220 active users, and 3:50 average engagement time, nearly 4 times more views and 6 times more active users than NeurIPS. Fully closed venues such as CVPR and ECCV record fewer than 35,000 views and under 1.5 minute average engagement (Yang, 2 Feb 2025).
The platform also reports comparative review dynamics. Review confidence statistics for 2024 are 3.53 ± 0.48 for ICLR, 3.58 ± 0.54 for NeurIPS, 3.54 ± 0.57 for ICML, and 3.64 ± 0.48 for CVPR. The paper states that ICLR shows a slightly lower concentration of high-confidence ratings among accepted papers than NeurIPS in 2024, consistent with more cautious evaluations in public settings. Discussion activity differs more sharply: ICLR exhibits broader and higher reply counts, with a maximum of 76 replies versus 49 for NeurIPS, and increasing medians and variance over years. Fully closed models are described as typically allowing only a one-time rebuttal, thereby limiting clarification and community input (Yang, 2 Feb 2025).
Evidence of grass-roots support is also reported through a survey on the Paper Copilot front page. The paper records over 228 responses across 20+ subfields and 50+ venues, with 57% of respondents willing to anonymously share their review scores from fully closed venues such as CVPR 2025 (Yang, 2 Feb 2025).
5. Strengths, limitations, and the case for regulated peer review
The position paper compares fully open, partially open, and fully closed review systems in explicitly balanced terms. Fully open systems are credited with real-time transparency, public visibility of reviews and author responses, richer iterative discussions, collective accountability, higher community engagement, and clearer auditability of irregularities. Their stated risks include reviewer caution or self-censorship under public visibility, subtle bias even within double-blind settings, exposure to public scrutiny or harassment, intellectual-property concerns for pre-decision disclosure, and potential plagiarism if rejected ideas remain visible (Yang, 2 Feb 2025).
Partially open systems are described as balancing transparency and privacy. Their strengths include post-decision public scrutiny and some accountability without exposing deliberation in real time. Their limitations include reduced opportunity for real-time correction and less iterative dialogue than fully open models. Fully closed systems are said to encourage frank private critique and protect confidentiality and proprietary ideas, which can align with patent and IP policies in some industry research groups; their disadvantages include limited transparency and accountability, fewer opportunities to detect inconsistencies such as authorship changes after acceptance, constrained author dialogue, and more difficulty enforcing AI or LLM usage policies in reviewing (Yang, 2 Feb 2025).
The paper’s central prescription is “regulated peer review.” This includes reviewer guidelines and training or mentorship, policies governing public critique and respectful discourse, anti-harassment safeguards, conflict-of-interest handling, explicit rules for AI/LLM usage in reviewing, and misconduct prevention processes that can detect and handle irregularities with public auditability. The paper also stresses balancing confidentiality with open science, especially for industry research involving patents or proprietary content (Yang, 2 Feb 2025).
Its recommendations extend beyond venue ideology to concrete operational steps. Venues are urged to choose a disclosure model and platform, adopt open APIs such as the OpenReview API, publish structured metadata on submissions, reviews, confidence, decisions, and timelines, create daily data snapshots for longitudinal analysis, publish public decision records and review artifacts, and enable community-driven supplements for closed phases or venues. Proposed success metrics include views, active users, session durations, click-through rate, discussion depth, confidence consistency, venue and year coverage, search rankings, and global reach (Yang, 2 Feb 2025).
6. Later developments and homonymous uses
A later Paper Copilot paper extends the project from a position paper and analytics platform into community-facing infrastructure and an open dataset for tracking the evolution of peer review in AI conferences. That work describes durable archives of time-varying review content, versioned snapshots, analytics at venue, institution, and country levels, and a large-scale empirical analysis of ICLR review dynamics. It also emphasizes reproducibility through released infrastructure and datasets, including GitHub-hosted paper lists and temporal OpenReview data (Yang et al., 15 Oct 2025).
The designation “Paper Copilot” has also been used for a different system in the academic-assistance literature. In that 2024 usage, Paper Copilot denotes a self-evolving, efficient LLM system for personalized academic assistance on top of a continuously updated corpus of arXiv papers. Its distinctive components are user profiles derived from publications, daily ingestion of new papers, and “thought-retrieval,” which retrieves accumulated trends, ideas, and prior answers in addition to raw documents. That system reports a reduction of average response time from 87.1s to 26.2s, or 69.92%, and a public demo hosted on Hugging Face (Lin et al., 2024).
Within the peer-review literature, however, Paper Copilot is chiefly associated with the argument that transparent review data should be treated as public research infrastructure. In that sense, the platform serves not only as a repository of descriptive statistics but as a mechanism for making review procedures measurable, contestable, and regulable within AI and machine learning research (Yang, 2 Feb 2025).