- The paper exposed a significant flaw in event-window studies of digital behavior, which misinterpret elevated post-event browsing as a causal response.
- Using empirical data and formal modeling, the authors identified two key factors causing the bias: endogenous time zero and episode-selection bias.
- The study demonstrated through empirical data on AI responses that actual user events preemptively elevate search activity, challenging conventional causal interpretations without further adjustments.
The problem: user-timed events as endogenous time zeros
Event-window studies of digital behavior routinely align users at moments the users themselves choose—opening an AI assistant, clicking a recommendation, visiting a product page—and interpret elevated activity afterward as an event effect. Iannelli and Ai show that this practice manufactures what they call endogenous time zero: because human activity is bursty, a user-timed focal event tends to fall inside an already-unfolding task episode, so the aligned curve can trace episode continuation rather than any response to the event. The behavioral mechanism is episode-selection bias—events are selected into latent high-intensity states that also drive outcomes—and it differs from classical activity bias [(Ventura et al., 2011)-style precedents; lewis2011here] in being within-user and time-varying: fixed effects and within-person matching compare the same person's busy moments to their quiet ones, which does not remove the confound.
The empirical signature appears directly in same-user, cross-surface web logs: around conversational-AI responses, browsing and search climb to a peak roughly eight minutes before time zero and decline smoothly through it, while placebo anchors stay flat. The selection is not specific to AI—first visits to shopping (3.6×), news (2.3×), coding/docs (5.3×), and reference (3.8×) sites are all preceded by broad browsing elevation over within-user placebos. A cross-surface lead–lag heatmap shows browsing and search running above placebo through entire ±90-minute windows with no onset at time zero, across all five focal domains.
The paper decomposes the naive contrast into an ATT plus an episode-selection term, proves episode-selection bias survives user fixed effects (Remark 1), and shows count/share divergence from one process (Example 1): when episodes raise all activity while the event locally substitutes for the outcome, raw counts rise even as shares fall. Proposition 1 establishes that the activity share is null-preserving only under a proportional-burst condition; Proposition 3 separates temporal admissibility (a past-only state estimate cannot introduce post-treatment bias) from adjustment validity (a noisy proxy generally fails conditional mean exchangeability).
The central theoretical result (Proposition 2) is a non-identification theorem: two generative models—one with a zero-effect event selected on a persistent latent state, one with conditional exchangeability and a strictly positive effect—induce the identical joint distribution over the entire single-surface observable process. The identified set for the ATT is (−∞,E[Y∣A=1]], unbounded below. This confines the paper's empirical strategy to falsification rather than estimation, and the authors say so explicitly.
The known-null experiment
The strongest evidence manufactures a case where the true answer is zero on real data. Pseudo-events—uniform-random timestamps passing the same pre-event landmark filter as real AI responses but causing nothing by construction—are matched to real events on strictly pre-event activity. Among washout-aligned, landmark-active responses (the 5.8% eligible subset of in-panel responses), real events show 4.32× [$3.34, 5.29$] post-event search lift over a within-user placebo; known-null timestamps show 3.42× [2.3×0]. The reproduced fraction is 2.3×1 [2.3×2], averaging over 40 analyst-drawn pseudo pools (design-conditional estimates 2.3×3–2.3×4). Sensitivity grids bound the share between 2.3×5 and 2.3×6, though the authors decline to call it robust.
The gradient by pre-event context is the sharpest finding: the reproduced fraction falls monotonically from 2.3×7 at detectably active anchors through 2.3×8 at sub-threshold ones to 2.3×9 at quiet ones, where the design detects no confounding. These are floors rather than decompositions—a synthetic calibration shows a noisy proxy recovers only 5.3×0–5.3×1 of a true fraction of 1—yet the quiet class's own lift remains substantial (5.3×2 against a null of 5.3×3), leaving two readings (real effects concentrate at quiet anchors, or episodes invisible to observed surfaces) that the design cannot separate.
A control "ladder" shows intermediate adjustments fail instructively: anchoring nulls on arbitrary page views alone manufactures a 5.3×4 association (an inspection-paradox effect), adding pre-event state matching raises it to 5.3×5–5.3×6, and even uniform-time anchors with full matching leave 5.3×7 against the real 5.3×8. Real events' apparent lift also declines monotonically with strictly past-defined episode age (5.3×9 to 3.8×0) while matched nulls show no comparable gradient.
Simulation and public benchmarks
In a zero-effect two-state Markov burst simulation, six common estimators—user fixed effects, activity matching, pre-activity stratification, recent-activity intensity summaries, negative-control movement—land between 3.8×1 and 3.8×2 against a true 3.8×3, indistinguishable from the naive 3.8×4. Only an oracle on the true burst state recovers the null; a two-sided smoothed HMM removes 3.8×5 (admissible only when the event cannot move the recovery stream); a genuinely past-only filtered forecast removes 3.8×6. Pre/post differencing flips sign under asymmetric burst placement (3.8×7), showing that where an event falls inside its episode sets the apparent sign of the "effect."
Public plasmodes complete the benchmark: on MovieLens, a naive window returns 3.8×8 an injected effect while a both-sides intensity smoother (admissible there only by construction) recovers the truth; on Wikipedia daily views, a past-only Poisson-HMM forecast removes 3.8×9 of the excess at eight states. A discriminant check on mechanically collected Wikipedia spikes validates the pre-event-elevation diagnostic at scale (AUC ±900, anticipated versus surprise timing).
What changes the answer
The paper organizes responses into four tiers: falsification tests (pre-event trajectories, negative controls, active-window placebos, pseudo-events), alternative estimands (activity share, episode-level comparisons, first-observed events), bias-reduction methods (episode matching, inverse-intensity weighting, past-only latent-state filtering, cross-surface proxies), and identification (exogenous timing, instruments, verified proxy conditions). Cross-surface leave-one-out indices attenuate ±901–±902 of the naive AI→search excess (±903 down to ±904), read as attenuation rather than bias removed.
The recommended default—an episode-level comparison—is itself audited: among completed episodes, search share is higher with AI present (±905 to ±906 points), but a presence placebo (news pageviews) shifts share three to four times more, so the surplus reads as residual selection. A twelve-item diagnostic protocol packages into burstcheck, which flags five of six checks on its bundled zero-effect demo. A worked case applies the audit to a stylized industry "shopping appetite" finding; every line moves, and only compositional and discrete estimands survive.
Limitations and open questions
The authors are candid about scope. The headline result describes a narrow population—washout-isolated, detectably active anchors comprising ±907 of in-panel AI responses—and is not extrapolated to quieter or unobserved episodes. Panel terms prevent reporting sample sizes, replaced with influence statistics; dropping the ten most influential users moves the headline fraction to ±908, outside the design-conditional range, and the influence tail is heavier than the coverage study's synthetic data can rule out undercovering for. The nondifferential-proxy assumption behind Proposition 3(b)'s sign guarantee is violated by the panel's own landmark eligibility rule, though simulations suggest the conclusion survives. Adjustments reduce excess association without identifying effects; confidence intervals do not propagate state-model estimation error; and the breakdown-value analysis (Appendix on identification bounds) concedes that in a constructed zero-effect world the residual posterior gap exceeds the computed frontier. Two prevalence audits—one of arXiv papers, one of industry publications—failed informatively and support no prevalence claim in either direction. The most direct open question the paper names is testing the covariate-instrument condition of Freyaldenhoven et al.—that a cross-surface index responds to episode state while remaining unaffected by the focal event—which would convert these diagnostics into an estimator.
Conclusion
A user-timed event is not automatically an exogenous time zero. On real behavioral trajectories, timestamps with true effect exactly zero reproduce most of the apparent post-event lift at detectably active anchors, and no functional of single-surface event-aligned data identifies the effect without further assumptions. The practical shift the paper argues for is in the unit of comparison—from the event to the episode—and in the standard of evidence, with diagnostics determining whether a naive contrast has earned a causal reading and design-based variation required to establish whatever remains.