- The paper demonstrates that TRIBE’s group-averaged fMRI model fails to correlate with moment-level YouTube rewatch data.
- The study employs GFP reduction, supervised regression, and position controls to isolate and test neural predictors against baseline audio-visual metrics.
- The findings imply that off-the-shelf predicted-fMRI signals cannot be used for fine-grained engagement forecasting, highlighting the need for subcortical-inclusive approaches.
Summary of "A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps" (2607.01400)
Introduction and Background
Recent advances in multimodal brain-encoding architectures, exemplified by TRIBE—the 2025 Algonauts challenge winner—have enabled high-fidelity prediction of cortical fMRI responses to naturalistic video, leveraging fused visual (V-JEPA 2), auditory (Wav2Vec-BERT), and textual (Llama-3.2) representations. Simultaneously, neuroforecasting studies have established links between measured neural signals (e.g., fMRI, EEG) and aggregate population behavior, especially in cultural and economic domains. This work investigates whether predicted neural signals from such models can forecast moment-level behavioral engagement, operationalized here as YouTube's "most replayed" heatmaps—a public, aggregate signal of user-initiated replay behavior at the per-second scale.
Methodology
A total of 48 diverse YouTube videos with available replay heatmaps across 11 content categories were processed. TRIBE—as released, with group-averaged subject embeddings—predicted per-second cortical responses for each video. The high-dimensional cortical prediction was reduced via the global field power (GFP), representing overall cortical drive, yielding a per-second engagement curve. This engagement curve was correlated, after quadratic detrending to account for position-related confounds, with the video’s "most replayed" curve to determine whether neural activity fluctuations tracked viewer replay behavior.
Low-level audio (loudness) and visual (frame motion) baselines were included for comparison. Additional supervised probes were conducted: ridge regression on the principal components of the predicted cortex and the foundation input features (visual, auditory, textual) to directly predict replay signals under matched/mismatched and strengthened position controls. The study further examined region-specific readouts (functional networks and ROIs), controlled for autocorrelation, and tested video-level engagement prediction.
Key Findings
Null Prediction of Replay Heatmaps
The primary finding is the absence of meaningful correlation between TRIBE-predicted engagement and YouTube replay heatmaps. The pooled position-controlled partial correlation is +0.058 (95% CI [−0.04, 0.15]; p=0.23), statistically indistinct from zero and not exceeding baseline correlations from loudness (+0.04, p=0.74). Category-level analyses showed inconsistent, non-systematic predictions across content types.
Robust Bounds on Effects
An equivalence test restricts the possible true effect size to below r≈0.14, and a Bayes factor of 3.2 provides moderate evidence for the null over the alternative. The reliability of the behavioral target is high (split-half ρ≈0.82), precluding label noise as a confound.
Failure of Cortical and Regional Probes
Regionally restricted readouts (visual, auditory, parietal, salience, frontal, vmPFC, anterior insula, ACC) yielded partial correlations at or near zero, with no functional network recovering a significant relationship to replay heatmaps. This null result persisted under non-parametric (circular-shifted) permutation controls.
Supervised Probes and Position Control
Supervised, cross-validated regression probes, which initially showed moderate raw correlations (up to r=0.47 across videos), were revealed as artifacts of dominant, shared temporal structure in the heatmap rather than content-specific predictive power. When position controls were applied more stringently via cubic spline detrending or matched/mismatched evaluation, the apparent signal collapsed (r≈0.14, p=0.87 for cortex), showing no video-specific prediction.
Source of Lost Signal
The same supervised probe applied to foundation feature streams indicated a weak, borderline video-specific signal in the visual stream alone (matched/mismatched p=0.004–$0.06$), with no evidence from audio, text, or predicted cortex. These results localize any residual content information upstream of the fMRI encoding process. The group-averaged nature of TRIBE explicitly collapses idiosyncratic, content-specific variance toward the population mean.
No Video-Level Engagement Prediction
Mean or peak TRIBE engagement signals did not correlate with video-level engagement metrics (view count, like count; all Spearman ∣ρ∣<0.28), likely reflecting the restricted range of already-popular videos required for heatmap availability.
Inter-Subject Correlation and Generalization
Custom per-subject encoders trained on naturalistic fMRI data recovered standard fMRI predictivity (in-domain p=0.740, cross-domain p=0.741) but failed to track replay behavior via predicted ISC (p=0.742, p=0.743). This explicitly tested and ruled out transfer of one of the best-established neuroforecasting effects at the time scale of interest.
Implications
These results demonstrate that high-accuracy, group-averaged fMRI encoding models are not, by default, predictive of within-item behavioral engagement at fine temporal scales as reflected in large-scale human rewatch behavior. The implication is twofold:
- Theoretical: Group-based modeling averages out the idiosyncratic variation—across subjects and across stimuli—that is likely essential for predicting behavioral outcomes derived from subcortical or individually variable cortical sources. The canonical neuroforecasting effect is largely subcortical (e.g., NAcc, vmPFC), which is invisible to surface-only models like fsaverage5-based TRIBE.
- Practical: Off-the-shelf use of deep predicted-fMRI signals or common derived features (GFP, ISC, or mean/peak cortical drive) for engagement prediction in media, marketing, or content optimization tasks is unsupported. Caution is essential, especially given the ability of simple temporal trends to confound apparent predictivity. Domain-specific, supervised models trained on raw behavioral or explicit neural targets remain necessary.
Limitations and Future Directions
Principal limitations include the exclusion of subcortical (reward-related) ROIs, the use of most-replayed as a behavioral proxy with possible biases, analysis restricted to a 60-second window of already-popular videos, and relatively small p=0.744. Future work should investigate subcortical-inclusive encoding, per-subject reliability readouts, improved engagement targets (e.g., creator-side audience-retention), and larger, more balanced datasets. The released pipeline and manifest enable these investigations.
Conclusion
Predicted fMRI drive from group-averaged, multimodal encoding models like TRIBE does not forecast moment-level YouTube rewatch behavior beyond trivial baselines. This null result is robust, bounded, and mechanistically localized to both stimulus and model architecture limitations. Only a faint, borderline content-specific signal persists in visual input upstream of the cortex, and no evidence supports the use of predicted cortex signals for engagement forecasting at this granularity. Whether subcortical modeling or improved behavioral targets will change this outcome remains to be determined.