Papers
Topics
Authors
Recent
Search
2000 character limit reached

Event-Time Confounding Under Bursty Human Dynamics

Published 21 Aug 2026 in cs.HC and stat.ME | (2608.21294v1)

Abstract: Studies of digital behavior often align users at moments they choose, such as opening an AI assistant, clicking a recommendation, or visiting a product page, and interpret higher activity afterward as an event effect. We show how this creates an endogenous time zero: the event occurs during an ongoing task episode, so the aligned curve can trace episode continuation rather than a response to the event. In same-user, cross-surface web logs, AI, shopping, news, coding, and reference events are all preceded by broad activity increases that peak before time zero. Our strongest test uses known-null timestamps that cause nothing. Among the 5.8% of AI responses meeting strict pre-event activity and washout criteria, these timestamps show 3.42 times the post-event search activity of a within-user placebo, compared with 4.32 times for real events. The fraction of excess reproduced by the known null falls from 0.56 at detectably active moments to -0.04 at quiet moments, where the design detects none. We formalize this episode-selection bias, prove that a single-surface event window cannot separate it from a genuine effect without additional assumptions, and show in zero-effect simulations why user fixed effects and coarse activity matching can fail: the confound is within-user and time-varying. We provide a diagnostic protocol, public-data benchmarks, and burstcheck, a lightweight audit tool. User-timed events may have real effects, but post-event volume does not identify them by default; studies should compare similar episodes with and without the event.

Authors (2)

Summary

  • The paper exposed a significant flaw in event-window studies of digital behavior, which misinterpret elevated post-event browsing as a causal response.
  • Using empirical data and formal modeling, the authors identified two key factors causing the bias: endogenous time zero and episode-selection bias.
  • The study demonstrated through empirical data on AI responses that actual user events preemptively elevate search activity, challenging conventional causal interpretations without further adjustments.

The problem: user-timed events as endogenous time zeros

Event-window studies of digital behavior routinely align users at moments the users themselves choose—opening an AI assistant, clicking a recommendation, visiting a product page—and interpret elevated activity afterward as an event effect. Iannelli and Ai show that this practice manufactures what they call endogenous time zero: because human activity is bursty, a user-timed focal event tends to fall inside an already-unfolding task episode, so the aligned curve can trace episode continuation rather than any response to the event. The behavioral mechanism is episode-selection bias—events are selected into latent high-intensity states that also drive outcomes—and it differs from classical activity bias [(Ventura et al., 2011)-style precedents; lewis2011here] in being within-user and time-varying: fixed effects and within-person matching compare the same person's busy moments to their quiet ones, which does not remove the confound.

The empirical signature appears directly in same-user, cross-surface web logs: around conversational-AI responses, browsing and search climb to a peak roughly eight minutes before time zero and decline smoothly through it, while placebo anchors stay flat. The selection is not specific to AI—first visits to shopping (3.6×3.6\times), news (2.3×2.3\times), coding/docs (5.3×5.3\times), and reference (3.8×3.8\times) sites are all preceded by broad browsing elevation over within-user placebos. A cross-surface lead–lag heatmap shows browsing and search running above placebo through entire ±90\pm90-minute windows with no onset at time zero, across all five focal domains.

Formalization and non-identification

The paper decomposes the naive contrast into an ATT plus an episode-selection term, proves episode-selection bias survives user fixed effects (Remark 1), and shows count/share divergence from one process (Example 1): when episodes raise all activity while the event locally substitutes for the outcome, raw counts rise even as shares fall. Proposition 1 establishes that the activity share is null-preserving only under a proportional-burst condition; Proposition 3 separates temporal admissibility (a past-only state estimate cannot introduce post-treatment bias) from adjustment validity (a noisy proxy generally fails conditional mean exchangeability).

The central theoretical result (Proposition 2) is a non-identification theorem: two generative models—one with a zero-effect event selected on a persistent latent state, one with conditional exchangeability and a strictly positive effect—induce the identical joint distribution over the entire single-surface observable process. The identified set for the ATT is (,E[YA=1]](-\infty,\, E[Y\mid A{=}1]], unbounded below. This confines the paper's empirical strategy to falsification rather than estimation, and the authors say so explicitly.

The known-null experiment

The strongest evidence manufactures a case where the true answer is zero on real data. Pseudo-events—uniform-random timestamps passing the same pre-event landmark filter as real AI responses but causing nothing by construction—are matched to real events on strictly pre-event activity. Among washout-aligned, landmark-active responses (the 5.8%5.8\% eligible subset of in-panel responses), real events show 4.32×4.32\times [$3.34, 5.29$] post-event search lift over a within-user placebo; known-null timestamps show 3.42×3.42\times [2.3×2.3\times0]. The reproduced fraction is 2.3×2.3\times1 [2.3×2.3\times2], averaging over 40 analyst-drawn pseudo pools (design-conditional estimates 2.3×2.3\times3–2.3×2.3\times4). Sensitivity grids bound the share between 2.3×2.3\times5 and 2.3×2.3\times6, though the authors decline to call it robust.

The gradient by pre-event context is the sharpest finding: the reproduced fraction falls monotonically from 2.3×2.3\times7 at detectably active anchors through 2.3×2.3\times8 at sub-threshold ones to 2.3×2.3\times9 at quiet ones, where the design detects no confounding. These are floors rather than decompositions—a synthetic calibration shows a noisy proxy recovers only 5.3×5.3\times0–5.3×5.3\times1 of a true fraction of 1—yet the quiet class's own lift remains substantial (5.3×5.3\times2 against a null of 5.3×5.3\times3), leaving two readings (real effects concentrate at quiet anchors, or episodes invisible to observed surfaces) that the design cannot separate.

A control "ladder" shows intermediate adjustments fail instructively: anchoring nulls on arbitrary page views alone manufactures a 5.3×5.3\times4 association (an inspection-paradox effect), adding pre-event state matching raises it to 5.3×5.3\times5–5.3×5.3\times6, and even uniform-time anchors with full matching leave 5.3×5.3\times7 against the real 5.3×5.3\times8. Real events' apparent lift also declines monotonically with strictly past-defined episode age (5.3×5.3\times9 to 3.8×3.8\times0) while matched nulls show no comparable gradient.

Simulation and public benchmarks

In a zero-effect two-state Markov burst simulation, six common estimators—user fixed effects, activity matching, pre-activity stratification, recent-activity intensity summaries, negative-control movement—land between 3.8×3.8\times1 and 3.8×3.8\times2 against a true 3.8×3.8\times3, indistinguishable from the naive 3.8×3.8\times4. Only an oracle on the true burst state recovers the null; a two-sided smoothed HMM removes 3.8×3.8\times5 (admissible only when the event cannot move the recovery stream); a genuinely past-only filtered forecast removes 3.8×3.8\times6. Pre/post differencing flips sign under asymmetric burst placement (3.8×3.8\times7), showing that where an event falls inside its episode sets the apparent sign of the "effect."

Public plasmodes complete the benchmark: on MovieLens, a naive window returns 3.8×3.8\times8 an injected effect while a both-sides intensity smoother (admissible there only by construction) recovers the truth; on Wikipedia daily views, a past-only Poisson-HMM forecast removes 3.8×3.8\times9 of the excess at eight states. A discriminant check on mechanically collected Wikipedia spikes validates the pre-event-elevation diagnostic at scale (AUC ±90\pm900, anticipated versus surprise timing).

What changes the answer

The paper organizes responses into four tiers: falsification tests (pre-event trajectories, negative controls, active-window placebos, pseudo-events), alternative estimands (activity share, episode-level comparisons, first-observed events), bias-reduction methods (episode matching, inverse-intensity weighting, past-only latent-state filtering, cross-surface proxies), and identification (exogenous timing, instruments, verified proxy conditions). Cross-surface leave-one-out indices attenuate ±90\pm901–±90\pm902 of the naive AI→search excess (±90\pm903 down to ±90\pm904), read as attenuation rather than bias removed.

The recommended default—an episode-level comparison—is itself audited: among completed episodes, search share is higher with AI present (±90\pm905 to ±90\pm906 points), but a presence placebo (news pageviews) shifts share three to four times more, so the surplus reads as residual selection. A twelve-item diagnostic protocol packages into burstcheck, which flags five of six checks on its bundled zero-effect demo. A worked case applies the audit to a stylized industry "shopping appetite" finding; every line moves, and only compositional and discrete estimands survive.

Limitations and open questions

The authors are candid about scope. The headline result describes a narrow population—washout-isolated, detectably active anchors comprising ±90\pm907 of in-panel AI responses—and is not extrapolated to quieter or unobserved episodes. Panel terms prevent reporting sample sizes, replaced with influence statistics; dropping the ten most influential users moves the headline fraction to ±90\pm908, outside the design-conditional range, and the influence tail is heavier than the coverage study's synthetic data can rule out undercovering for. The nondifferential-proxy assumption behind Proposition 3(b)'s sign guarantee is violated by the panel's own landmark eligibility rule, though simulations suggest the conclusion survives. Adjustments reduce excess association without identifying effects; confidence intervals do not propagate state-model estimation error; and the breakdown-value analysis (Appendix on identification bounds) concedes that in a constructed zero-effect world the residual posterior gap exceeds the computed frontier. Two prevalence audits—one of arXiv papers, one of industry publications—failed informatively and support no prevalence claim in either direction. The most direct open question the paper names is testing the covariate-instrument condition of Freyaldenhoven et al.—that a cross-surface index responds to episode state while remaining unaffected by the focal event—which would convert these diagnostics into an estimator.

Conclusion

A user-timed event is not automatically an exogenous time zero. On real behavioral trajectories, timestamps with true effect exactly zero reproduce most of the apparent post-event lift at detectably active anchors, and no functional of single-surface event-aligned data identifies the effect without further assumptions. The practical shift the paper argues for is in the unit of comparison—from the event to the episode—and in the standard of evidence, with diagnostics determining whether a naive contrast has earned a causal reading and design-based variation required to establish whatever remains.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.