- The paper validates Gemini 2.5 Flash with eight sampled video frames and text as the strongest scalable auditor, reaching Cohen’s κ=0.42 while showing that text-only analysis is inadequate.
- The audit of 36,971 videos finds Italy had the highest baseline harm exposure, including a 48.6% passive-scrolling rate for age-19 personas, driven mainly by sexually suggestive and Italy-exclusive content.
- TikTok keyword search raised harmful-content exposure to 35–56%, sharply narrowing age differences and delivering restricted content to minors without visible search friction, although exposure returned to baseline afterward.
Overview and research questions
Saffari and Pierri (Politecnico di Milano) present a sockpuppet audit of TikTok across three EU countries—France, Italy, and Sweden—and four age personas (13, 16, 19, 40), using multimodal LLMs (MLLMs) as automated harm annotators. The study addresses three questions: (RQ1) which MLLM configuration agrees with native-speaker annotators sufficiently to run at scale; (RQ2) what is the marginal value of sampled frames versus native video versus text-only input; and (RQ3) how does harm exposure vary by age persona, country, and platform signal (passive For-You-page scrolling vs. active keyword search). The corpus comprises 36,971 unique videos collected via VPN-routed, locale-matched accounts driven by a Tampermonkey userscript, with the full audit costing approximately $49 in API spend.
Data collection
The passive phase (30 December 2025–11 January 2026) records pure FYP scrolling; the active phase (13–19 April 2026) implements a within-account scroll-pre → SEARCH → scroll-post cycle using 21 native-language keywords spanning seven probed harm categories. The country selection is motivated by DSA transparency data on per-language moderator allocation (France ~620, Italy ~396, Sweden ~98 moderators). A notable descriptive finding: English dominates passive feeds everywhere (30–48% native-language share only 10.5–20.6%), while native-language search queries lift native content to 30–39%. The authors are appropriately careful that a higher harm rate under harm-keyword search is expected by construction; the substantive findings concern magnitude, age-gradient collapse, and absence of visible moderation friction.
MLLM validation (Stage 1)
Four models (Gemini 2.5 Flash, Qwen3-VL-32B, GPT-4o-mini, Mistral Large 3) were evaluated under three input conditions against a 300-video reference set labeled independently by two native-speaker annotators per country with joint resolution. Gemini 2.5 Flash with eight sampled frames plus text (E3) wins at aggregate Cohen's κ=0.42, roughly half the per-call cost of native-video upload—and no configuration reaches moderate agreement (κ≥0.41) in every country. Agreement improves monotonically with visual input for Gemini (E1 κ=0.17 → E2 $0.38$ → E3 $0.42$); all non-Gemini configurations sit below κ=0.30, and text-only input fails outright (no model clears fair agreement), consistent with prior evidence that video LLMs exploit language priors rather than genuine temporal reasoning.
Two caveats deserve emphasis. First, per-country κ for Gemini-E3 tracks pre-resolution inter-annotator κ almost proportionally—the model reaches ~69–74% of the human-consensus ceiling in each country—suggesting the cross-country variation reflects reference-label quality, not model capability. Second, strict primary-subcategory agreement among jointly-harmful items is only 29% (54% under primary-or-secondary matching), so category-level prevalence estimates rest on weak ground. Notably, Gemini's confusion profile shows it under-flags harm (48 false negatives vs. 24 false positives), inverting the usual over-moderation concern.
Cross-country and cross-age prevalence (Stage 2)
Applying Gemini-E3 to a phase-stratified 10% sample (~10,500 verdicts), Italy carries the highest harm rate at every age persona, with Italian age-16 at 40.4% overall and Italian age-19 reaching 48.6% under passive scrolling—the audit's highest rate—against ~22.5% in France at the same age. Sexually suggestive content dominates flagged items in every cell. Three decompositions localize the Italian lead: it is concentrated in one category (Sexually Suggestive contributes 23.8 pp of Italy's 39.0% passive rate), it is carried entirely by Italy-exclusive content (shared-circulation videos flag at rates indistinguishable from FR/SW), and it survives a language control (English-language videos served to Italian accounts flag at 32.4% vs. 19.5% in France). The authors conclude the pattern points to the country-specific content pool rather than translation artifacts or annotator thresholds, while conceding that supply-side vs. moderation-capacity explanations cannot be distinguished from outside.
An important nuance: a per-country precision/recall recalibration derived from Stage-1 leaves the ordering intact at ages 13–19 but inverts the age-40 comparison (corrected FR-40 ≈49% overtakes IT-40 ≈38%), which the authors report as a caveat on the strongest reading.
Search-phase amplification
The central exposure finding: the SEARCH endpoint returns 35–56% harmful content, a 1.5–7.5× lift over scroll-pre in ten of twelve combinations, with no visible search-time block or warning. Two structural results follow:
- The spike is temporary: scroll-post reverts to within a few percentage points of scroll-pre in all twelve combinations (account-clustered hierarchical bootstrap confirms this), so keyword search elevates exposure during the session without persistent feed drift.
- The age gradient collapses under search: SEARCH harm rates compress into a 35–56% band across all twelve combinations, with 13-year-old personas receiving 24–35 pp of 18+-restricted content—within a few points of adults in the same countries. However, the design cannot distinguish age-insensitive retrieval from keyword-pool overlap across personas, since identical keywords were issued by every persona; an item-overlap analysis would discriminate between these readings.
A policy-tier decomposition shows the flat total-rate age gradient mixes a rising restricted tier (consistent with age-gating intent) with a non-declining universally-prohibited tier for minors—so the result is not uniform age-blindness.
Provider blocks as systematic measurement bias
Gemini's safety layer refuses ~1.1% of Stage-2 inputs, but these refusals cluster non-randomly: Nudity (~4.8%) and Sexually Suggestive (~2.6%) show the highest per-category block rates among SEARCH-attributable refusals. Reported prevalences are therefore under-estimates precisely for the categories the audit targets; a worst-case sensitivity bound shifts headline rates by at most 1–3 pp and preserves both the IT > SW ≥ FR ordering and the Italian ceiling effect.
Limitations
The authors are explicit about several constraints. The Stage-1 κ=0.42 was measured on 300 videos stratified by country and age but not phase, so transfer to the full Stage-2 distribution is unverified and phase-dependent auditor error cannot be ruled out. At scale, native-video (E2) reports harm 3–12 pp higher than E3 consistently; which modality better matches human judgment on Stage-2 items cannot be resolved without a second annotation pass. Reference labels come from guideline-grounded graduate-student annotators, not platform enforcement truth. Collection spans a single time window; sockpuppet personas lack real-user behavioral richness, making findings upper-bound claims about algorithmic reach rather than point estimates of real exposure; VPN-routed accounts may have been handled differently by anti-abuse systems, an unmeasured residual.
Conclusion
The paper establishes two substantive results: harm-keyword search—not recommender amplification—is the dominant exposure pathway in this audit, returning 35–56% harmful content with no visible search-time friction while reverting immediately afterward; and baseline algorithmic exposure is sharply country-dependent, with Italy most exposed at every age through a category-concentrated, pool-specific mechanism. Methodologically, it demonstrates that a validated MLLM auditor can scale cross-national youth-safety audits at API costs near $50, with annotation throughput, not compute cost, as the binding constraint—but the moderate ceiling of model-human agreement (κ=0.42) means MLLM auditing complements, rather than replaces, human annotation on policy-edge judgments.