- The paper shows that post-review AI feedback encouraged 78.4% of surveyed reviewers to consider revisions, but only 56.9% actually revised, usually by adding content or improving clarity rather than changing judgments.
- Reviewers valued feedback that improved tone, specificity, and actionability, yet often dismissed it as generic or shallow because it focused on surface writing instead of novelty, rigor, and contribution.
- The study finds that AI feedback did not reduce reviewers’ sense of ownership or accountability, while more experienced reviewers were less likely to revise and participants favored earlier, transparent, human-controlled AI support.
Context and motivation
Peer review operates under mounting strain: CHI submissions more than doubled from 3,182 in 2023 to 6,730 in 2026, and ICLR submissions grew from 3,422 in 2022 to 11,672 in 2025, while the volunteer reviewer base has not scaled correspondingly. Venues have diverged sharply in their response—CVPR prohibited LLM use in any part of review writing in 2025, whereas ICLR 2025 became the first top-tier computer science conference to deploy an official AI feedback tool within its review workflow. That tool, a multi-LLM pipeline built on Claude Sonnet 3.5 (two actor models, an aggregator, a critic, and a formatter), scanned submitted reviews for vagueness, possible misunderstandings of the paper, and unprofessional tone, then emailed reviewers targeted, optional suggestions for improving clarity and actionability. The tool did not alter reviews or influence decisions.
Prior work on AI support for reviewing had relied on hypothetical scenarios or Wizard-of-Oz prototypes. This study by Chen, Zhong, Brumby, and Cox (CHI '26) provides the first empirical account of how reviewers experienced such a tool in a live, high-stakes review process, complementing the ICLR organizers' own randomized trial of more than 20,000 reviews, which reported aggregate behavioural effects (roughly one-quarter of reviews revised) but not reviewers' lived experiences (Thakkar et al., 13 Apr 2025).
Method
The authors conducted a mixed-methods study combining an online survey (N = 51 eligible ICLR 2025 reviewers, drawn from 92 responses) with semi-structured interviews with nine survey participants. The survey, structured around the HALIE framework, used 7-point Likert items covering perceived usefulness, feedback quality, impact on decision-making, ownership, and reuse intentions, analysed with one-sample t-tests against the neutral midpoint. Interviews followed Braun and Clarke's reflexive thematic analysis. The sample skewed male (88.2%) and early-career (43.1% aged 25–34; 37.3% PhD students), a limitation the authors acknowledge explicitly.
Survey findings: perception and behaviour
Perceptions were lukewarm but not hostile. Two items differed significantly from neutral, both positively: perceived relevance (M=4.51, p=0.032, d=0.31) and willingness to use such feedback in the future (M=4.49, p=0.031, d=0.31). Intention to revise was the strongest effect (M=4.59, p=0.012, d=0.37), with 78.4% of participants reporting they considered revising. By contrast, ratings of constructiveness (M=3.88), actionability (p=0.0320), and usefulness (p=0.0321) did not differ from neutrality, and 27.5% agreed the feedback was at times inappropriate (p=0.0322, p=0.0323). Notably, there was no evidence that the tool reduced reviewers' sense of ownership (p=0.0324) or accountability (p=0.0325); reviewers continued to regard the review as their own.
A central quantitative finding is the intention–action gap: while 78.4% considered revising, only 56.9% actually did so. Among those who revised, changes clustered in content addition (9 of 19), clarity improvement (6), and actionability (5)—elaboration and polish rather than substantive re-evaluation. Among non-revisers, the dominant reason was perceived low value: the feedback was described as generic, redundant, or focused on expression rather than content. The only significant demographic predictor of actual revision was review experience (p=0.0326, p=0.0327): more experienced reviewers were less likely to edit their reviews, suggesting confidence in their initial judgements.
Interview findings: benefits, drawbacks, and envisioned futures
Interviews surfaced four benefits: polishing and extending reviews, prompting reflection, improving tone and professionalism, and making feedback more actionable. Drawbacks were more consistently reported. Reviewers found the suggestions overly general and shallow ("the feedback from the AI is very generic, nothing valuable"), checklist-like ("add more references"), and consequently easy to disregard. A distinct resistance pattern emerged around the tool's positioning: some reviewers read the post-hoc critique as second-guessing completed work—"I don't like this AI. It's like a critic: I've already finished my draft"—and one participant predicted senior reviewers would be still less receptive.
Participants' visions for future AI support divided into capability enhancement and fairness safeguards. They imagined AI as a collaborative reasoner synthesizing paper and review content, a provider of intellectual labour relief (summarization, fact-checking, citation-graph retrieval), a bridge for cross-domain understanding and dialogue, and—at the system level—a quality safeguard for area chairs that could fact-check claims, flag low-effort or internally inconsistent reviews, and normalize scoring. Across all visions, participants insisted evaluative judgement remain human, and they called for transparency mechanisms such as disclosing the percentage of AI contribution to a review.
The paradox of perception
The paper's central analytical claim is a paradox of perception: reviewers often acknowledged that AI feedback improved clarity and specificity yet still judged it unhelpful, because the tool targeted surface expression—polishing, elaboration, tone—while reviewers locate the value of their work in intellectual judgement about novelty, rigour, and contribution. Observable improvement did not translate into perceived value. Timing amplified this: because feedback arrived only after submission, reviewers felt their intellectual work was already done, making revisions appear costly. The authors also note a divergence between their sample (nearly half revised) and the ICLR deployment at large (much lower uptake), attributing it partly to self-selection of more engaged reviewers—an honest concession that qualifies the generalizability of their revision figures.
A second conceptual contribution concerns role reversal. In most human–AI scholarly workflows, humans prompt and AI generates; here the AI issued prompts and reviewers responded. The tool functioned as a "meta-commentator" that did not change review decisions but reshaped rhetorical norms—nudging reviewers toward greater elaboration and, in the authors' framing, prompting the community to reconsider whether longer, more detailed reviews are actually "better" reviews, or simply more burdensome ones.
Limitations
The authors are explicit about three constraints. First, the sample was skewed toward male, early-career researchers and may suffer self-selection bias, limiting claims about senior reviewers' experiences. Second, the study relied on self-reported perceptions and intentions rather than independent measures of review quality; ethical constraints precluded examining confidential reviews. Third, findings are specific to ICLR 2025's post-hoc, discussion-based OpenReview model, and may not transfer to venues with single review–rebuttal cycles or to earlier, opt-in, or real-time intervention designs. These caveats bear directly on the design implications: recommendations for early-stage, verifiable support (summarization, reference checking, retrieval-augmented related-work surfacing) and mid-draft scaffolding remain hypotheses awaiting validation in live deployments.
Conclusion
This study documents the first empirical account of reviewer experiences with officially deployed post-review AI feedback in a high-stakes conference review process. Its key results—a significant intention–action gap, the negative correlation between review experience and revision, the paradox in which measurable improvement coexists with perceived uselessness, and preserved senses of ownership and accountability—indicate that the success of such tools depends less on their ability to improve review text than on their alignment with reviewers' professional identity and the timing of their interventions. The authors frame AI in peer review as simultaneously a design problem and a governance problem, arguing that venue-led systems, transparent AI involvement, and community deliberation over what "quality" means are prerequisites for sustainable integration. The open questions the paper leaves—how intervention timing shapes uptake, how senior reviewers respond, and whether rhetorical nudges at scale reshape review norms—are specific and empirically tractable.