Papers
Topics
Authors
Recent
Search
2000 character limit reached

Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis

Published 20 Aug 2026 in cs.AI and q-bio.NC | (2608.19902v1)

Abstract: AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher's computational environment under rules for admissible analyses, required checks and claim scope. In benchmarks, Brain Researcher increased first-choice tool-selection accuracy across seven models by 70.2 percentage points (23.3% without it versus 93.6% with it) and verifiable grounding from 4.6% to 22.0%. In collaborator-led and self-evolving studies, multiverse analyses exposed analytic-choice sensitivity, and scientific review classified claims as accepted, qualified, revised, blocked, rejected or deferred. By linking decisions to evidence and provenance, Brain Researcher embeds methodological judgment within the workflow, not after it.

Summary

  • The paper introduces Brain Researcher, a researcher-governed platform that links commitments, evidence, analyses, execution, and claim status through a provenance graph, sealed commitment cards, and auditable claim cards.
  • Across 60 tool-calling tasks and seven frontier models, the platform improved correct top-1 tool routing from 23.3% to 93.6% and verified evidence grounding from 4.6% to 22.0%, while still leaving most citations unverified.
  • The paper shows that multiverse analyses and explicit review rules can block unsupported claims and freeze follow-up studies, but a sign-blind review error demonstrates that expert inspection and independent replication remain essential.

Motivation and design rationale

The paper addresses a gap between executing analyses and establishing claims. Tool-using AI agents can now orchestrate neuroimaging pipelines, but successful command execution does not guarantee scientific validity: agents may select inappropriate tools, stop prematurely, optimize misaligned criteria, or reproduce the questionable research practices documented in human-executed science. The problem is acute in neuroimaging, where analytic flexibility is extreme — in the widely cited many-analysts study of a single fMRI dataset, no two teams used identical workflows and conclusions diverged substantially (2608.19902). Existing open-science infrastructure (BIDS, fMRIPrep, Nipype, BIDS Stats Models, FitLins) standardizes procedures and their alternatives but does not bind a selected route to the evidence and methodological conditions required for the claim it supports.

Brain Researcher is a researcher-governed, domain-specific agentic harness that runs inside the researcher's existing computational environment. Its central design principle is to preserve rather than replace scientific judgment: researchers specify prospectively what counts as an admissible analysis, which checks must pass, and the scope of any resulting claim; everything else is recorded for inspection. The system comprises a version-pinned tool registry with machine-readable specifications and pre-execution rule checking; a provenance-linked knowledge graph (BR-KG; 745,949 nodes, 2,461,469 edges) aligned to the OpenNeuro Vocabulary; an MCP server mediating model actions; and a review layer that assigns each claim one of six states (accepted, qualified, revised, blocked, rejected, deferred) under explicit adjudication rules.

Auditable claim records

The unit of work is a persistent episode linking question, evidence, admissible analyses, committed plan, execution, review, and condition-tagged conclusion. Two dated artifacts anchor this record: a commitment card, sealed with a content hash before any analysis runs, fixing the question, allowed alternatives, and success/failure criteria; and a claim card, written afterward by the review layer, recording claim state, scope, and checks passed or failed. A reviewer can audit the record without re-executing the analysis. This design makes methodological decisions visible at each stage while leaving them formally with the researcher — the system records and enforces checks but does not adjudicate judgments that resist formalization.

Benchmark results on tool selection and grounding

A paired benchmark across seven frontier models (Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, GLM-5.1, DeepSeek-V4-Pro, Kimi K2.5, Qwen3.6-Plus) compared performance with and without Brain Researcher's registry, knowledge graph, and constraint layer over 60 tool-calling tasks:

Metric Without BR With BR
Correct route/tool@1 23.3% 93.6%
Capability@1 49.8% 94.5%
Handoff score@1 47.4% 76.1%

All seven models improved on all three metrics (7/7 positive paired differences; exact two-sided Wilcoxon signed-rank p=0.016p=0.016 each), with mean gains of 70.2, 44.7, and 28.7 percentage points respectively. Condition-blind LLM judges credited any response reaching required capabilities through equivalent executable routes, so gains reflect capability attainment rather than naming Brain Researcher's tools. A routing ablation without direct KG calls still selected acceptable top-1 routes in 86.2% of 420 episodes.

A separate 50-question evidence-citation benchmark measured verified groundedness (cited evidence locatable and judged supportive): 4.6% without versus 22.0% with Brain Researcher — a 4.8-fold increase, though most rows still failed verification. The authors state plainly that grounding improved substantially without being solved. Among non-verified with-BR rows, 65% were judged real-but-off-topic and 28% partial support; none had fabricated or malformed majorities. A human audit of 20% of scored items agreed with 96% of judge verdicts (κ=0.94\kappa = 0.94), with discrepancies limited to one-step severity differences.

Two caveats bear directly on these numbers. The without-BR condition removes registry, knowledge graph, and constraint layer together, so the contrast measures the harness as a whole rather than isolating components. And because reference routes were curated with a co-author who does not develop the system, target construction may share vocabulary with the registry — a potential inflation of the routing gain.

Multiverse analyses expose claim sensitivity

Three collaborator-led studies tested whether multiverse analysis plus explicit constraints make claim sensitivity and status visible.

Schizophrenia NeuroMark audit (FBIRN cohort, N=363N=363; 5,460 edges per subject). Brain Researcher expanded three pre-specified hypotheses into a 480-specification multiverse. None was supported uniformly; all were recorded as qualified. NM-H2 (between-domain > within-domain group differences) showed a complete estimator split after sign-aware rescoring: 100% favorable under Pearson and Spearman, 0% under partial correlation and mutual information — an estimator-regime-dependent finding rather than a robust effect, with the underlying mechanism unresolved. NM-H1 and NM-H3 were weak (median Δ\DeltaAUC =0.032=-0.032 favoring edges; only 26.0% favoring between-domain loading mass).

This episode also supplied the paper's most consequential negative finding about its own review layer: after a server-side fault triggered fallback to a general-purpose coding agent, that agent scored any specification with permutation p<0.05p<0.05 as favorable regardless of sign, inflating apparent support for a directional hypothesis to near-universal levels. Automated review missed the error; a human reviewer caught it by inspecting code, outputs, and specification curves. Two checks were added (a directionality test and a warning on general-purpose-agent fallback), converting a one-off correction into an enforced check — but the failure demonstrates that formal review does not make expert inspection unnecessary.

Cocaine-use-disorder connectivity (SUDMEX CONN, N=138N=138). A 36-specification multiverse rejected all five pre-specified connectivity–behavior associations under SDMA-GLS (all Z<1.24Z<1.24, FDR q>0.58q>0.58); an exploratory screen over 70 combinations surfaced no FDR-surviving effect. The system blocked confirmatory promotion and converted the null into a replication plan — an example of prespecified checks withholding status from a result that would otherwise read as exploratory.

Cross-cultural social-cognition meta-analysis. Subgroup ALE on 21 studies produced an mPFC-topology interpretation, but the system blocked it as exploratory because subgroups held only k=6k=6–8 entries (below the recommended κ=0.94\kappa = 0.940), paradigm composition was imbalanced, and centroid shifts cannot establish non-overlapping distributions. The case ended in a paradigm-matched follow-up with no settled claim.

Across episodes, multiverse analysis exposed which findings depended on analytic choices, and review determined what each result could support. One structural limitation applies here: collaborators' hypotheses were pre-specified in their own protocols rather than sealed as commitment cards, so the NeuroMark record is a post-hoc audit rather than prospective governance.

Self-evolving episodes: frozen successor analyses

Two extended episodes tested whether intermediate evidence can redirect research within researcher-defined action spaces while successor analyses are frozen before execution.

HCP connectivity-based prediction. Starting from a published benchmarking study, Brain Researcher evaluated 116 candidate prediction pipelines for Cognition prediction (best discovery score κ=0.94\kappa = 0.941, whole-band coherence with ridge regression). After a selector audit, the researcher froze a coherence-based workflow; across 10 repeated nested-CV splits it achieved median κ=0.94\kappa = 0.942 versus κ=0.94\kappa = 0.943 for a matched reconstruction of the published procedure (median κ=0.94\kappa = 0.944; conditional one-sided κ=0.94\kappa = 0.945), higher in all 10 runs. Refit to four additional behavioral outcomes without retuning, it exceeded the reference in 37 of 40 comparisons (47 of 50 directional wins across all five outcomes). However, median out-of-sample κ=0.94\kappa = 0.946 was positive for only two variables, and multiplicity-aware transfer inference remained inconclusive — the workflow comparison is more robust than the predictive claims themselves.

TRIBE v2 representational geometry. Screening category contrasts by change in held-out discrimination, Brain Researcher set aside the unstable top-ranked lead (tools–voice) and proposed speech–tools: later layers bring the categories closer while preserving representational direction. The hypothesis and analysis were frozen before evaluation on new stimuli. Across three successive 48-item panels, normalized separation decreased in 11 of 12 collection-by-panel comparisons, all meeting the prespecified directional criterion — a direction-preserving contraction. Extension to four previously unused collections showed the same geometry in three, but the corrected collection-level test remained inconclusive (Holm-adjusted κ=0.94\kappa = 0.947).

In separately initiated sessions without Brain Researcher, the same coding agent completed substantial analyses but never generated and froze a follow-up study. Because these sessions were not matched controls, the authors correctly characterize this contrast as descriptive rather than causal.

Limitations and open questions

The paper concedes several limitations at points where they matter. Foundation-model and retrieval priors reflect well-represented literature and datasets, potentially biasing toward operationalized questions over negative results and low-resource populations. Several episodes rely on same-dataset multiverse or internal validation, which improves auditability but does not constitute independent replication. Public datasets (HCP, OpenNeuro) may have been encountered during model training; memorization cannot be ruled out, so results should not be treated as novel detections. Evidence-grounding judges included models also among the seven evaluated, a potential self-preference source, though the human audit mitigates this. Review-layer error was estimated against a 60-case calibration library assembled after the NeuroMark sign-direction check — zero false-accepts out of 16 invalid cases (rule-of-three 95% upper bound 19%) and zero false-blocks out of 5 valid controls — but this measures internal consistency on canonical scenarios, not field-scale escape rates of flawed claims. Runtime and researcher effort were unmeasured. Whether claim records function as shared, contestable infrastructure across laboratories remains an open objective requiring labeled real-world analyses.

Conclusion

Brain Researcher reframes the unit of AI-assisted research as a governed trajectory rather than a model output: commitments are sealed before results are observed, evidence carries provenance, multiverse analysis exposes specification sensitivity, and every claim receives an explicit, auditable state. The quantitative benchmarks show large harness-level gains in tool routing (23.3% → 93.6%) and meaningful but incomplete gains in evidence grounding (4.6% → 22.0%). The research episodes demonstrate both the value of this governance — blocked confirmatory promotion of unsupported effects, frozen successor analyses carried forward — and its limits, most clearly when automated review missed a sign-blind scoring error that only human inspection caught. The platform makes methodological conditions visible and enforceable; it does not substitute for expert judgment or independent replication, and field-scale validation of its review layer remains unresolved.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.