- The paper introduces a task-level benchmark that evaluates four commercial LLM systems across screening, extraction, analysis, and synthesis using 244 heterogeneous documents and expert-validated standards.
- The results show that no system performs best across every task: Claude led screening accuracy at 82.8%, GPT-5 achieved 91.8% recall, and synthesis quality varied by nuance, detail, and source type.
- The paper demonstrates a human-supervised routing workflow that combines shared retrieval with task-specific model selection, while identifying persistent weaknesses in limitations, future directions, visualization, worker-centered concerns, and underrepresented populations.
The Knowledge Synthesis Review (KSR) framework addresses a methodological gap in LLM-assisted evidence synthesis: most existing evaluations test single stages of a review, typically title-and-abstract screening, on homogeneous clinical or biomedical corpora. Shafqat, Patterson, and Liss instead decompose evidence synthesis into four cognitive tasks—screening, extraction, analysis, and synthesis—benchmark four commercial LLM-based systems against expert gold standards over deliberately heterogeneous sources, and use the resulting task-level performance profiles to route each task to the best-performing system under continuous human validation (2608.12741). The demonstration domain is the relationship between AI and labor markets, chosen because its evidence base is fragmented across peer-reviewed research, industry reports, policy briefs, and media commentary that differ sharply in quality, incentives, and reception.
Motivation and problem framing
The authors motivate the framework with a concrete asymmetry in how AI labor-market evidence circulates. A McKinsey projection that generative AI could automate up to 30% of work hours by 2030 received broad public visibility, while a 2025 Danish study of roughly 25,000 workers across 7,000 workplaces found no statistically significant first-year changes in aggregate output, hours, or wages—and reached primarily academic audiences. Because media framing demonstrably shapes public understanding and governance of AI, uneven reception of this kind is itself an argument for systematic multi-source synthesis at a pace traditional reviews, which take six months to over a year, cannot sustain.
Two research questions structure the study: RQ1 asks how widely used LLM systems perform across the four core synthesis tasks and heterogeneous source types against expert-annotated gold standards; RQ2 asks how task-level performance differences can be operationalized as a transparent orchestration workflow and what applying it at scale reveals about the capabilities and limits of LLM-assisted synthesis.
Corpus construction and benchmark design
Phase I assembled a 1,893-document corpus spanning 2020–2025: 324 research papers (RPs), 30 industry reports (IRs), 30 policy briefs (PBs), and 1,509 news/media/blog sources (NMBs). The imbalance toward NMB sources was preserved deliberately to reflect the actual evidence environment; the small IR and PB counts reflect genuine scarcity of such documents rather than sampling choices. From this corpus, a 244-document benchmark subset was drawn (103 RP, 30 IR, 30 PB, 81 NMB), deliberately over-representing the scarce, higher-credibility source types so each type received adequate evaluation.
Gold standards were established by two expert reviewers—one author-team member, one independent external evaluator—who labeled all 244 documents for screening prior to reconciliation. Inter-rater reliability was high: 92.2% raw agreement with Cohen's κ=0.804, stable across source types (κ between 0.718 for policy briefs and 0.831 for industry reports). Disagreements were resolved by consensus, yielding 183 retained documents for downstream tasks. All model outputs were anonymized and randomized before evaluation to limit system-specific bias. Notably, reliability statistics were computed only for screening; the analysis and synthesis rubric scores were reconciled through consensus without inter-rater statistics, a limitation the authors acknowledge.
Task protocols and evaluation metrics
Each task had distinct criteria. Screening used binary include/exclude labels evaluated via confusion matrices, precision, recall, specificity, accuracy, and F1. The protocol required iterative refinement: early prompts relying on article links failed because most models could not consistently access external URLs, and some models judged from titles alone despite richer JSON inputs. The final protocol used simplified natural-language instructions, JSON-formatted inputs, and standardized spreadsheet outputs. Extraction scored field-level accuracy against verified records, with exact string matching for fixed metadata fields and PERFECT/SEMI/NONE grading for open fields. Analysis used six dimensions (relevance, specificity, completeness, evidence, clarity and structure, plus a field-specific criterion) on a five-point rubric. Synthesis assessed coherence, coverage, cross-source integration, redundancy reduction, and handling of contradictions across six analytical dimensions.
The central empirical finding is that no system led on all tasks, and performance varied systematically by both task and source type:
| System |
Screening strength |
Extraction profile |
Analysis profile |
Synthesis profile |
| Claude Sonnet 4 |
Highest accuracy (82.8%), precision up to 70% on PB/NMB |
Most consistent overall; 76% match rate linking NMB articles to scholarly references |
Highest overall; 4.18 on RP methodology summaries |
Schematic summaries |
| GPT-5 |
Highest recall (91.8%) but specificity collapsed to 53.3% on PB |
Strong generally; sharp underperformance on PB titles/sources (32%) |
Strong methodology and findings (~4.0) |
Most nuanced outputs |
| Gemini 2.5 Pro |
Conservative; higher false negatives |
Near-perfect PB titles/sources |
Mid-range |
Schematic summaries |
| NotebookLM |
Conservative |
Mid-range, led in no field |
Scored as low as 1.53 on PB future directions |
Rich detail |
Extraction exceeded 90% agreement for titles and publishing sources across models but degraded substantially on author attribution and reference identification, with up to 40% of attempts yielding partial matches or failures. Analysis performance dropped markedly on forward-looking and critical-appraisal tasks: future-directions scores averaged only 2.9–3.5 out of 5, and limitations were similarly underdeveloped. In synthesis, GPT-5 and NotebookLM produced the most nuanced integrative narratives while all models preserved pluralism across contradictory sources, yet shared weaknesses in visualization, cultural diversity, and worker-centered concerns. The pattern—that fluent output does not guarantee complete or faithful coverage—is consistent with known failure modes of abstractive summarization and long-context information use.
Contamination check
Because the corpus consists of public documents potentially present in training data, the authors ran a contamination check comparing field-level extraction accuracy on post-cutoff documents (published after April 2025, the most recent cutoff among the three API-accessible systems) against full-corpus accuracy, restricted to NMB sources where 27 post-cutoff documents were available. Post-cutoff totals were 86.7% (GPT-5), 88.9% (Claude), and 85.2% (Gemini), versus full-corpus values of 80.0%, 80.3%, and 74.5%. No decline appeared on unseen documents, supporting the conclusion that prior exposure did not inflate results. The authors correctly note the confound: post-cutoff articles come largely from uniformly structured major outlets, so the higher absolute scores likely reflect source composition rather than capability. NotebookLM was excluded from this check because it performs internal retrieval and exposes no comparable cutoff.
Orchestration and scaled application
Phase III routes each task to its best-performing system over a shared retrieval-augmented generation (RAG) layer built on ChromaDB, ensuring all models operate over identical evidence and reducing dependence on parametric knowledge. Outputs are returned in harmonized formats (.xlsx, JSON, narrative files), with expert review validating routed outputs throughout. An important caveat: the routing strategy was motivated by benchmark differences but was not validated end-to-end against a single-system baseline on held-out documents, and no time or cost accounting quantifies the efficiency claim—the authors flag both as immediate priorities.
Applied to the full corpus, the routed workflow surfaced cross-source asymmetries invisible to single-source synthesis. Media oscillated between alarmist and optimistic framings; industry reports emphasized productivity and business opportunity while underweighting algorithmic surveillance, hiring bias, and safety nets; academic findings remained discipline-fragmented. All four source types agreed that AI simultaneously displaces and creates jobs, with routine low-skill roles facing highest automation risk, but only one source type treated reskilling as a strong current market priority—a gap between acknowledged necessity and present employment-system emphasis. The synthesis also identified systematic blind spots: small and medium enterprises, worker well-being (including surveillance stress and diminished autonomy), informal economies, gender disparities, and Global South contexts receive insufficient attention relative to productivity and competitiveness themes. These observations are presented as outputs of evidence synthesis—patterns, asymmetries, and gaps—not as direct empirical estimates of AI's labor-market effects.
Limitations and open questions
The authors concede several constraints plainly. The corpus is English-language and text-only, leaving multilingual, visual, and non-textual evidence untested. The evaluated models represent a 2025 baseline whose relative rankings may shift with new releases. The routed workflow lacks validation against single-system and human-only baselines with explicit time and cost accounting. Reliability statistics are absent for analysis and synthesis rubric scores, which remain sensitive to reviewer expertise and rubric design. The contamination check covers only NMB extraction, not other tasks or source types. Open questions include whether task-level routing generalizes across domains beyond AI and labor markets, and whether larger reviewer panels change rubric-score stability.
Conclusion
The KSR framework contributes a replicable, model-agnostic design for governing LLM assistance in evidence synthesis: task decomposition, evaluation rubrics, and routing logic persist while the routing table is re-derived as systems evolve. Its empirical core demonstrates that LLM performance in synthesis is task-specific and source-specific rather than universally reliable—no system dominated, interpretive and cross-source tasks remained weakest, and expert judgment remained essential precisely where automation would be most valuable. The framework's value proposition is augmentation rather than automation: computational scale for screening and structured extraction, with human accountability governing interpretation, disagreement preservation, and final claims.