Papers
Topics
Authors
Recent
Search
2000 character limit reached

WhatsApp Tiplines: Encrypted Fact-Checking

Updated 7 July 2026
  • WhatsApp tiplines are opt-in, end‑to‑end encrypted channels where users submit texts, images, and videos for fact-checking, serving as a crowd-sourced intake for suspicious content.
  • They leverage multilingual and multimodal data processing with advanced clustering and human verification to analyze misinformation during elections and crises.
  • Operational workflows span from content intake and triage to verification and feedback, providing privacy-preserving, real-time fact-checking and data for research.

WhatsApp tiplines are opt-in channels that fact-checking organizations run on WhatsApp so users can forward suspicious content for verification on an end‑to‑end encrypted platform. In research on misinformation during elections and crises, they are treated as citizen-facing intake systems that capture text, images, video, audio, and links that would otherwise remain largely inaccessible to external observation on encrypted services. Across studies of India’s 2019 and 2021 elections and Brazil’s 2022 general election, WhatsApp tiplines appear as a privacy-preserving crowdsourcing mechanism, a workflow for human fact-checking, and a data source for multilingual and multimodal claim analysis (Kazemi et al., 2021, Shahi et al., 22 Jul 2025, Hale et al., 2024).

1. Definition and institutional role

A tipline is a dedicated service on WhatsApp, typically a phone number or WhatsApp account, to which users forward messages they want fact-checked. Users submit text, images, links, videos, and audio, and operators reply to some content with a fact-check or trusted guidance. In the Indian 2021 election study, users sent “tips” to advertised WhatsApp numbers, and a tip was treated as a “claim” if it contained fact-check-worthy content (Shahi et al., 22 Jul 2025). In the 2019 Indian election study, the “Checkpoint” WhatsApp tipline was led by PROTO and Pop-Up Newsroom, a joint project of Meedan and Fathm, with technical assistance from WhatsApp, and Meedan’s open-source software was used to operate the tipline (Kazemi et al., 2021). In Brazil’s 2022 election, three fact-checking organizations operated WhatsApp tiplines, complemented by a national WhatsApp chatbot run by the Tribunal Superior Eleitoral (TSE) (Hale et al., 2024).

The institutional role of tiplines is dual. Operationally, they provide a bidirectional channel through which citizens surface potentially misleading content and receive verification or guidance. Analytically, they provide a lens into misinformation circulating in private chats and groups despite end-to-end encryption. The 2019 Indian case study explicitly argues that a crowd-sourced tipline is a useful source for discovering content to fact-check and that it captures WhatsApp-native viral content better than open-platform-driven fact-check pipelines (Kazemi et al., 2021). The Brazilian study similarly situates tiplines as channels for surfacing suspected falsehoods already circulating in private spaces, while the TSE bot captured procedural concerns and integrity questions (Hale et al., 2024).

This suggests that the distinctive value of a tipline is not merely messaging access, but access coupled to user intent: submissions are already pre-filtered by suspicion, urgency, or uncertainty. A plausible implication is that this makes tipline streams operationally different from passive monitoring of public groups or open platforms.

2. Workflow and verdicting

In practice, tipline organizations collect tips, triage them, perform verification, publish a fact-check article, and notify the user with the result. The 2021 Indian election study describes manual verification against credible sources such as government portals, academics, and medical doctors, followed by assignment of a verdict (Shahi et al., 22 Jul 2025). If a claim cannot be verified or is outside scope, it is labeled inconclusive or out of scope. When a fact-check already exists, some tiplines can return a response immediately; the paper shows a demo tipline with instant feedback for previously checked questions (Shahi et al., 22 Jul 2025).

Verdict systems vary at source level and may be normalized analytically. In the 2021 study, 12 original verdict labels were normalized to four: false, partially false, true, and other, with “other” bundling out of scope, inconclusive, and “please send more” (Shahi et al., 22 Jul 2025). The study explicitly did not re-evaluate content; it only normalized labels. In the Brazilian study, fact-checkers triaged greetings and spam versus actionable claims, clustered duplicates, and then undertook verification, with responses delivered via the same channels (Hale et al., 2024).

A workflow-oriented representation of the recurring pipeline across the studies is shown below.

Stage Description Examples/tools named
Intake Users forward suspicious content to WhatsApp numbers or bots Text, images, video, audio, links
Triage Remove spam/greetings, de-duplicate, route by language/topic Clustering, language ID
Verification Human fact-checking against credible sources Government portals, academics, medical doctors
Verdicting Assign or normalize labels false, partially false, true, other
Feedback Notify users and publish fact-checks Tipline reply, fact-check article

The data also support a distinction between immediate and delayed response. Previously checked items may permit instant feedback, whereas novel claims generally require human investigation. In the 2021 Indian study, fact-checkers generally required a couple of days to debunk a new claim and share the result with users (Shahi et al., 22 Jul 2025).

3. Data regimes, scale, and temporal dynamics

The empirical scale of tipline operations differs sharply across contexts. In India’s 2021 assembly elections, the study began with 4,945 textual tips, selected 950 textual claims that were fact-checked, and analyzed a final sample of 580 unique claims from 451 users in English, Hindi, and Telugu during March 1–May 31, 2021 (Shahi et al., 22 Jul 2025). These data were gathered across four states and one union territory amid the Delta wave of COVID‑19. English contributed 271 claims, Hindi 199, and Telugu 110, with mean claim lengths of 131, 160, and 179 words respectively, and an overall mean of 150 words (Shahi et al., 22 Jul 2025).

In contrast, the 2019 Indian national election tipline received 157,995 total messages over March 1–June 30, 2019, yielding 82,676 unique items, including 37,823 unique text messages, 10,198 unique links, and 34,655 unique images (Kazemi et al., 2021). The 2022 Brazilian election study reports 49,422 submissions from 14,959 unique users to three fact-checker tiplines during 2022-09-01 to 2022-11-15, plus 223,621 submissions to the TSE bot’s fact-checking feature (Hale et al., 2024).

These studies also document event-sensitive temporal dynamics. In India 2021, tip volume rose in March, peaked in mid‑April during final voting phases, dipped at the end of April, and increased in English in mid‑May (Shahi et al., 22 Jul 2025). In Brazil 2022, tipline submissions peaked on the first-round and run-off election days, with another peak immediately after the run-off driven by inquiries about results; the number of new clusters per day tracked submissions closely (Hale et al., 2024).

The 2021 Indian study provides a compact topic and verdict profile for its 580-claim corpus.

Measure Value
COVID‑19 claims 215 (37%)
Election claims 143 (25%)
Other claims 222 (38%)
False 105
Partially false 33
True 58
Other verdicts 384

The relation between topic mix and verdict mix is important. The study notes that approximately 66% of tips were out of scope or lacked sufficient information, which is consistent with the size of the “other” verdict category (Shahi et al., 22 Jul 2025). This is not the same as the topical “other” category: one concerns subject matter, the other verdict outcome.

4. Multilingual and multimodal analysis

A central research contribution of the recent literature is the treatment of WhatsApp tiplines as multilingual and multimodal corpora rather than merely text channels. The 2021 Indian election study focuses on English and Hindi as high-resource languages and Telugu as a low-resource language (Shahi et al., 22 Jul 2025). Language identification used pycld3, and claims with confidence below 0.90 were manually reviewed by a language expert; 36 tips had their languages corrected (Shahi et al., 22 Jul 2025). The preprocessing pipeline removed numbers and URLs, merged Hindi and Hindi‑English code-switching into Hindi, translated Hindi and Telugu claims into English via Google Translate, and manually checked translations. Code-switching in Hinglish was present and merged into Hindi; transliterated content was translated (Shahi et al., 22 Jul 2025).

For semantic comparison, the study used a BERT-based sentence transformer, Indian XLM‑R, to embed claims and computed cosine similarity,

cos(x,y)=xyxy\cos(x,y) = \frac{x \cdot y}{\lVert x \rVert \, \lVert y \rVert}

followed by hierarchical agglomerative clustering with SciPy’s ward linkage (Shahi et al., 22 Jul 2025). The procedure initialized each claim as its own cluster, constructed a condensed distance matrix using cosine distance, and merged clusters bottom-up to a single root. Specific hyperparameters and external validation metrics were not reported. Frequent word analysis was frequency-based rather than TF‑IDF-based, though TF‑IDF is discussed in the background (Shahi et al., 22 Jul 2025).

The Brazilian 2022 study extends this picture to a fully multimodal claim-matching setting. Text clustering used paraphrase-multilingual-MPNet-base-v2 sentence embeddings, cosine similarity, and single-link hierarchical clustering with threshold τ=0.875\tau = 0.875 (Hale et al., 2024). Preprocessing removed URLs, lower-cased text, removed punctuation, and replaced accents and diacritics with closest ASCII equivalents to reduce orthographic mismatch across platforms (Hale et al., 2024). For images, the study used PDQ perceptual hashing and normalized Hamming-distance similarity defined as s=1ds = 1 - d, with threshold $0.7$ (Hale et al., 2024). For videos, it used TMK embeddings and tmk-clusterize with thresholds $0.7$ for both level‑1 and level‑2 (Hale et al., 2024).

The 2019 Indian study similarly combined multilingual sentence embeddings for text with cosine similarity thresholded at $0.9$, online single-link hierarchical clustering, FAISS for efficient retrieval, and PDQ plus DBSCAN for image clustering (Kazemi et al., 2021). Reported performance for the text matching model on similar data was F1=0.73F1 = 0.73 overall and approximately $0.85$ for English and Hindi (Kazemi et al., 2021).

Across these studies, multilingual sentence embeddings and perceptual hashing serve different but complementary functions. Embeddings group semantically similar claims expressed in different languages or paraphrases; perceptual hashes and temporal video embeddings detect near-duplicate media variants. This suggests that effective tipline analysis requires a hybrid representational stack rather than a single universal similarity model.

5. User behavior, overlap, and audience structure

Tipline usage exhibits skewed participation and structured audience segmentation. In the 2021 Indian election study, 451 unique users submitted 580 claims, and 59 users sent multiple claims totaling 189 requests (Shahi et al., 22 Jul 2025). Cross-language overlap based on pseudonymous IDs was present but limited: 8 users submitted in both English and Hindi, 9 in English and Telugu, and 3 in Telugu and Hindi (Shahi et al., 22 Jul 2025). No user submitted fact-checked claims to multiple organizations, suggesting that each organization maintained a unique audience (Shahi et al., 22 Jul 2025).

The Brazilian study finds a similar pattern, but with measurable partial overlap across three fact-checker tiplines. Of 6,383 users who submitted at least two messages, 93% sent all messages to the same tipline, 7% to two tiplines, and 0.6% to all three (Hale et al., 2024). The study interprets this as evidence of distinct tipline audiences. User activity was also heavy-tailed: 54% of users submitted only one item, the most active user submitted 299 items, and 5% of users contributed 41% of unique items (Hale et al., 2024).

The Indian 2021 study also reports topical overlap among repeat users: 14 users submitted both other–election claims, 14 election–COVID‑19, 22 other–COVID‑19, and 8 users submitted claims in all three categories (Shahi et al., 22 Jul 2025). Because organizations shared only pseudonymous identifiers and not phone numbers or names, these overlap analyses were deliberately limited to aggregate patterns (Shahi et al., 22 Jul 2025).

A concise comparison of audience overlap findings is useful.

Study Within-language or within-tipline behavior Cross-organization behavior
India 2021 Limited cross-language user overlap None across organizations
Brazil 2022 93% of multi-message users stayed with one tipline 7% used two; 0.6% used all three
India 2019 Sustained user engagement noted Specific cross-organization comparison not reported

The absence or near-absence of cross-organization overlap does not imply informational isolation in a strong sense, but it does indicate that separate organizations may cultivate distinct user bases. A plausible implication is that consortium-based coordination can improve coverage without assuming automatic user migration across tiplines.

6. Findings on content, similarity, and operational performance

The 2021 Indian election study reports strong cross-lingual thematic commonality. On translated English text, word counts were 5,433 for English, 4,939 for Hindi, and 4,161 for Telugu. Shared vocabulary comprised 541 words for English–Telugu, 500 for English–Hindi, 420 for Hindi–Telugu, and 102 words shared across all three languages (Shahi et al., 22 Jul 2025). Shared words included COVID‑19 terms such as “ventilator,” election figures such as “mamata,” and general concerns such as “education,” “business,” and “smell” (Shahi et al., 22 Jul 2025). Clustering identified a top-level division between a mixed group involving gas and mineral water prices, mobile numbers, and fraud, and a group dominated by election and COVID‑19 content, including state COVID‑19 reports, vaccines to Canada, rally mask protocol breaches, scams, and local COVID‑19 news (Shahi et al., 22 Jul 2025).

The topical balance also varied by language. English claims were distributed as 118 COVID‑19, 35 election, and 118 other; Hindi as 56 COVID‑19, 67 election, and 76 other; Telugu as 41 COVID‑19, 41 election, and 28 other (Shahi et al., 22 Jul 2025). Average claim lengths by category were 198 words for election, 176 for COVID‑19, and 94 for other (Shahi et al., 22 Jul 2025). The study therefore concludes that COVID‑19 claims outnumbered election claims overall, while Hindi and Telugu carried comparatively more election content than English (Shahi et al., 22 Jul 2025).

Operationally, the same study measured the average time from tip receipt to fact-check publication at 2.92 days overall, with language-specific means of 2.15 days for English, 2.98 days for Hindi, and 3.2 days for Telugu (Shahi et al., 22 Jul 2025). It reports that 94% of claims were debunked within four days, though some required weeks (Shahi et al., 22 Jul 2025). Factors influencing turnaround by category or claim type were not quantified.

The 2019 Indian case study adds a different performance perspective: coverage and lead–lag relative to public WhatsApp groups and ShareChat. For images shared 100 or more times in public groups, 67% were also submitted to the tipline; for text messages shared 100 or more times, 23% appeared in the tipline (Kazemi et al., 2021). Of 257 annotated text claim clusters, 93% matched at least one message in public groups, with an average per-cluster match percentage of 91% (Kazemi et al., 2021). For the top 10% most shared images in public groups, about 80% were first submitted to the tipline, even though approximately 50% of all images appeared in public groups first (Kazemi et al., 2021).

The Brazilian 2022 study emphasizes novelty and platform divergence. It reports that 78% of clusters contained only one item and that the relationship between number of users and novel content was approximately linear, indicating no saturation: adding users continued to surface new claims (Hale et al., 2024). It also shows low overlap between fact-checker tiplines and the TSE bot—18% for videos, 1% for images, and less than 0.01% for text—while overlap counts for shared items were positively correlated across sources, with r0.67r \approx 0.67 for videos, $0.55$ for images, and τ=0.875\tau = 0.8750 for text (Hale et al., 2024).

Taken together, these findings indicate that tiplines simultaneously serve as channels for repeated, high-salience rumors and as discovery systems for novel claims. This dual role complicates optimization: a system must both collapse duplicates efficiently and remain sensitive to genuinely new narratives.

7. Limitations, ethics, and implementation trajectories

The literature is explicit about the methodological and ethical constraints of tipline research. Because WhatsApp is end-to-end encrypted, tipline datasets are opt-in and crowdsourced rather than representative samples of platform traffic (Kazemi et al., 2021, Hale et al., 2024). In the 2021 Indian study, only text tips were analyzed; images, audio, and video were excluded, and the final sample of 580 claims was modest (Shahi et al., 22 Jul 2025). Translation required manual correction, and some residual errors may remain. Language coverage was concentrated in English, Hindi, and Telugu, while other regional languages had negligible verified tips (Shahi et al., 22 Jul 2025). In Brazil, only timestamps and randomized IDs were available, the TSE bot user count was unknown due to anonymity, and the platform comparison was constrained by missing APIs and collection outages (Hale et al., 2024).

Ethical handling of submissions is treated as fundamental rather than ancillary. The 2021 Indian study reports that all analyzed data were pseudoanonymous; no phone numbers or personal identifiers were shared; users were informed of data use for fact-checking and non-commercial academic research; only aggregated analysis was conducted; and strict data access, storage, and auditing procedures ensured accountability and privacy (Shahi et al., 22 Jul 2025). The 2019 Indian study similarly states that all WhatsApp messages were anonymized and analyses were conducted at macro level with strict data safeguards (Kazemi et al., 2021). The Brazilian study notes that all WhatsApp data were anonymized prior to analysis, with phone numbers replaced by random IDs and only timestamps retained (Hale et al., 2024).

The implementation literature also links tiplines to privacy-preserving intervention architectures. The 2020 study on already-debunked misinformation proposes an on-device matching system in which WhatsApp maintains vetted digital fingerprints of previously labeled false content and distributes them to clients for local checking before or after encryption (Reis et al., 2020). In that proposal, tiplines act as sensors for emerging variants: user-submitted content, once fact-checked, enriches the debunk corpus and enables future on-device warnings or forwarding restrictions without server-side scanning (Reis et al., 2020). The empirical rationale is that a large share of misinformation-image dissemination in public groups occurred after fact-check publication—40.7% in Brazil and 82.2% in India, or 71.7% without one outlier in the Indian data (Reis et al., 2020).

Within the operational studies themselves, recommendations converge on multilingual staffing, strong triage, and duplicate detection. The 2021 Indian paper recommends multilingual teams including regional language experts, triaging and prioritization flows to handle surges and de-duplicate repeated claims, clear user feedback about turnaround time and supported languages, and cross-region collaboration to address nationally shared narratives across local languages (Shahi et al., 22 Jul 2025). It also suggests considering on-device, opt-in tools using hashes for images or text similarity to pre-alert users to previously debunked content as complementary to tiplines (Shahi et al., 22 Jul 2025). The Brazilian study, while marking many recommendations as inferential, points toward OCR and ASR integration, audio-aware video matching, hierarchical cross-platform matching, and UI designs that elicit enough detail for local claim disambiguation (Hale et al., 2024).

A plausible synthesis is that WhatsApp tiplines are evolving from narrow fact-check request channels into modular infrastructures for multilingual intake, multimodal claim normalization, human verification, and privacy-preserving downstream intervention. The current literature does not present them as a complete solution to misinformation on encrypted platforms. Rather, it presents them as a tractable and empirically productive interface between private messaging environments, professional fact-checking, and computational methods for claim discovery and matching.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to WhatsApp Tiplines.