Papers
Topics
Authors
Recent
Search
2000 character limit reached

Inverse Surprising Popularity

Updated 14 July 2026
  • Inverse Surprising Popularity (ISP) is a phenomenon where realized, predicted, and perceived support diverge, leading to a distorted signal of true prevalence.
  • The ranked-voting approach uses pairwise reduction with partial votes and ordinal predictions to recover true rankings despite noisy, inverse-popular signals.
  • Additional models show that network exposure bias and popularity-based ranking feedback can concentrate support, exemplified by the few-get-richer effect in social media settings.

Inverse Surprising Popularity (ISP) is not defined by name in the papers considered here. In this literature, the term is most naturally understood as referring to settings in which realized support, predicted support, and locally perceived support diverge systematically, so that naive popularity is a distorted signal of a latent global truth, ranking, or prevalence. That interpretation is explicit in surprisingly popular voting and its ranking-recovery extension, where respondents report both their own answer and a prediction of others’ answers (Hosseini et al., 2021). It is closely mirrored, though not formally identical, in directed-network models of prevalence distortion induced by the friendship paradox (Alipourfard et al., 2019) and in popularity-based ranking systems exhibiting the few-get-richer effect (Germano et al., 2019). Taken together, these works place ISP at the intersection of meta-belief aggregation, network exposure bias, and algorithmic visibility dynamics.

The most direct formal lineage for ISP is classical surprisingly popular voting. In that setting, each participant reports a vote and a prediction of other participants’ votes, and the aggregation rule selects the option whose actual support exceeds its predicted support in the relevant sense. The ranking paper restates the classical motivation succinctly: classical democratic aggregation works when the majority is relatively accurate, whereas surprisingly popular voting can recover the ground truth even when experts are in minority (Hosseini et al., 2021).

The three papers considered here do not describe a single unified ISP formalism. One paper develops a ranking-recovery analogue of classical surprisingly popular voting; one analyzes neighborhood-based prevalence distortion in directed networks; and one studies an algorithmic inverse-popularity effect under popularity-based ranking. A common source of confusion is to treat these as the same object. The literature instead separates at least three distinct mechanisms. First, there is the canonical vote–prediction discrepancy mechanism of surprisingly popular voting. Second, there is directed-network perception bias, where local exposure overweights structurally prominent actors. Third, there is ranking-feedback amplification, where popularity-updated visibility can cause a smaller class of items to attract more aggregate traffic than a larger class.

This distinction matters for ISP because the phrase “inverse popularity” can refer either to an aggregation rule that corrects naive vote counts by using meta-predictions, or to a structural process in which visibility and exposure make observed popularity systematically misleading. The ranking paper is the closest to ISP in protocol design; the directed-network and popularity-ranking papers are best understood as mechanistic analogies or network-structural foundations rather than direct implementations of the standard surprisingly popular method (Alipourfard et al., 2019, Germano et al., 2019).

The most explicit extension of surprisingly popular logic beyond single-option truth discovery is the ranked-voting framework of “Surprisingly Popular Voting Recovers Rankings, Surprisingly!” (Hosseini et al., 2021). Let AA be a set of mm alternatives, L(A)L(A) the set of rankings over AA, and πL(A)\pi^* \in L(A) the true ranking. Each voter ii observes a noisy ranking σiL(A)\sigma_i \in L(A), drawn from a signal distribution conditional on π\pi^*. The model uses the posterior

P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},

and from that derives another voter’s signal distribution,

P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).

If full votes and full predictive distributions over rankings were elicited, the original SP score is

mm0

Under the condition

mm1

the paper quotes the theorem from Prelec et al.: mm2

The practical difficulty is that direct SPV over rankings requires predictions over all mm3 rankings. The paper therefore studies partial votes and partial predictions. The vote formats are Top vote and Rank vote. The prediction formats are Top prediction and Rank prediction. This yields the four principal elicitation formats Top-Top, Top-Rank, Rank-Top, and Rank-Rank, together with the no-prediction baselines Top-None and Rank-None.

The methodological core is a pairwise reduction. For each pair mm4, the method extracts a binary pairwise vote mm5 and a pairwise prediction mm6. Because classical SPV requires cardinal predictions but the elicited predictions are ordinal, the paper introduces a two-parameter conversion using

mm7

These parameters are learned from training data. The pairwise SP-style scores are then

mm8

and

mm9

If L(A)L(A)0, the method records L(A)L(A)1; otherwise it records L(A)L(A)2.

Because pairwise decisions are made independently, the output can be a cyclic tournament rather than a globally coherent ranking. The paper therefore uses the resulting tournament in two ways: predicting the top alternative by choosing the alternative that defeats the largest number of others, and evaluating full-ranking recovery by Kendall Tau distance to the ground truth. The empirical study uses 720 MTurk participants and 7,200 responses over geography, movies, and paintings tasks with objective ground-truth rankings. The central findings are that prediction helps substantially, Rank-Rank performs best, and even a little prediction information helps surprisingly popular voting outperform classical approaches. Particularly notable is the result that Top-Rank significantly outperforms Rank-Top, suggesting that richer prediction elicitation can matter more than richer vote elicitation.

3. Directed-network prevalence distortion as an ISP-adjacent mechanism

“Friendship Paradox Biases Perceptions in Directed Networks” formalizes a different but closely related object: the gap between true global prevalence and locally perceived prevalence in a directed network (Alipourfard et al., 2019). The network is L(A)L(A)3, with L(A)L(A)4. A directed edge L(A)L(A)5 means that L(A)L(A)6 is a friend of L(A)L(A)7, or equivalently that L(A)L(A)8 follows L(A)L(A)9; the edge direction is the direction of information flow. For a node AA0, AA1 denotes out-degree, interpreted as number of followers, and AA2 denotes in-degree, interpreted as number of friends.

The paper defines three sampling-based random variables. A random node AA3 is uniform over nodes: AA4 A random friend AA5 is sampled proportional to out-degree: AA6 A random follower AA7 is sampled proportional to in-degree: AA8 Because total in-degree equals total out-degree,

AA9

Each node has a binary attribute πL(A)\pi^* \in L(A)0. The true global prevalence is πL(A)\pi^* \in L(A)1. The global feed-level perceived prevalence is πL(A)\pi^* \in L(A)2. For node πL(A)\pi^* \in L(A)3, the local perceived prevalence is

πL(A)\pi^* \in L(A)4

and the average local perception is πL(A)\pi^* \in L(A)5. This decomposition is especially important for ISP-style reasoning because it distinguishes actual prevalence, friend-weighted prevalence, and neighborhood-level prevalence.

A major result is the explicit bias formula

πL(A)\pi^* \in L(A)6

The paper also rewrites this bias in terms of the Pearson correlation πL(A)\pi^* \in L(A)7, out-degree standard deviation πL(A)\pi^* \in L(A)8, and binary-attribute standard deviation πL(A)\pi^* \in L(A)9. The interpretation is direct: if trait-holders have higher out-degree, they are overrepresented in others’ feeds, and degree heterogeneity strengthens the distortion.

The paper further defines local perception bias,

ii0

and introduces attention

ii1

It derives

ii2

Sufficient conditions for positive local perception bias are

ii3

and

ii4

The appendix states that these conditions imply

ii5

The same paper formalizes four directed variants of the friendship paradox. Two hold in any directed network: random friends have more followers than random nodes, and random followers have more friends than random nodes. Two additional variants require positive in/out-degree correlation: random friends have more friends than random nodes, and random followers have more followers than random nodes. A plausible implication for ISP is that “perceived popularity” is not a single object. The paper explicitly emphasizes that global and local biases can even have opposite signs.

4. Estimation from biased perceptions in directed networks

The same directed-network framework also provides a concrete estimation-from-biased-perceptions procedure through Follower Perception Polling (FPP) (Alipourfard et al., 2019). The goal is to estimate

ii6

the true global prevalence. FPP samples ii7 independent nodes from the follower distribution

ii8

and uses the estimator

ii9

The motivation is a bias–variance tradeoff. Random followers tend to have larger in-degree than random nodes, so they observe more friends, and their neighborhood averages σiL(A)\sigma_i \in L(A)0 have lower variance than the perceptions of random nodes. The paper’s bias theorem states

σiL(A)\sigma_i \in L(A)1

From the appendix,

σiL(A)\sigma_i \in L(A)2

so FPP intentionally accepts the same bias as friend-weighted prevalence rather than directly debiasing local perceptions.

The paper also gives an explicit unbiased modification,

σiL(A)\sigma_i \in L(A)3

Methodologically, this is the clearest inversion step in the directed-network setting: degree-weighted local information can be corrected by inverse-degree weighting.

For variance, let σiL(A)\sigma_i \in L(A)4 be the adjacency matrix, σiL(A)\sigma_i \in L(A)5, and σiL(A)\sigma_i \in L(A)6. The degree-discounted bibliographic coupling matrix is

σiL(A)\sigma_i \in L(A)7

If σiL(A)\sigma_i \in L(A)8 is connected and non-bipartite, the paper provides an exact variance expression and the upper bound

σiL(A)\sigma_i \in L(A)9

where π\pi^*0 and π\pi^*1 is the second largest eigenvalue of π\pi^*2.

The empirical validation uses Twitter data from 2014. Starting from 100 users active on California ballot initiatives in 2012, the collection expands to 5,599 seed users, then to over 600K users total, with over 18M hashtag mentions. Each hashtag π\pi^*3 is treated as a binary trait π\pi^*4 if user π\pi^*5 used hashtag π\pi^*6. Among the 1,153 most popular hashtags, each used by more than 1,000 people, 865 had positive local bias. For #ferguson, global prevalence is π\pi^*7 and local perceived prevalence is π\pi^*8, so it appeared about four times more popular than it actually was. In synthetic polling experiments on an induced subgraph of 5,409 users, FPP has lower variance than both Intent Polling and Node Perception Polling, and for sample budget π\pi^*9 it outperforms both in MSE for most hashtags. Even at P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},0, it outperforms IP in more than 80% of cases and NPP in more than 55% of cases.

5. Popularity-based ranking and the few-get-richer effect

“The few-get-richer: a surprising consequence of popularity-based rankings” analyzes a distinct inverse-popularity mechanism in search and recommender systems (Germano et al., 2019). There are P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},1 items partitioned into two classes, P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},2 and P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},3. Sequentially arriving users have heterogeneous class preferences. For user P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},4, the item-level propensity absent ranking is

P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},5

This assumption is central: class-level preference is divided uniformly across items within a class, so fewer items imply more concentrated per-item support.

Popularity-based ranking updates according to cumulative clicks. With attention bias parameter P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},6, the click probability is

P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},7

Aggregate class-level traffic share is

P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},8

The paper’s central result is the few-get-richer effect: the fewer the items in a given class, the higher the total share of traffic that class may collectively attract. The analytical thresholds are sharp. If

P(πσi)=P(σiπ)P(π)πL(A)P(σiπ)P(π),P(\pi^* \mid \sigma_i) = \frac{P(\sigma_i \mid \pi^*) \cdot P(\pi^*)}{\sum_{\pi' \in L(A)} P(\sigma_i \mid \pi') \cdot P(\pi')},9

a stable limit ranking places all class-1 items at the top. If

P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).0

class 1 is sufficiently large that class 0 is the small class, and the P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).1 items are bottom-ranked. The formal proposition compares two environments P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).2 and P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).3 differing only in P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).4 and P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).5, and states that there exists P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).6 such that for any P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).7, total clicking probability on class 1 is strictly greater in the environment with fewer class-1 items, provided P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).8 is sufficiently small: P(σjσi)=πL(A)P(σjπ)P(πσi).P(\sigma_j \mid \sigma_i) = \sum_{\pi^* \in L(A)} P(\sigma_j \mid \pi^*) \cdot P(\pi^* \mid \sigma_i).9

Mechanistically, concentrated niche demand lifts a small number of minority-class items in the ranking; once they reach high positions, indifferent users and rank-biased users reinforce them. The paper’s simulations examine mm00, mm01, uniform initialization, and varying mm02, mm03, and mm04. The reported pattern is that CTR to mm05 decreases as mm06 increases, and stronger attention bias strengthens the effect.

The online experiment uses 786 Amazon Mechanical Turk participants. Participants identify as cat person, dog person, or neither, then choose among 20 clickable cat or dog pictures. There are four dynamic conditions with popularity-updated ranking and four static matched controls. In the dynamic conditions, dog pictures receive more than 50% of traffic in all cases: D1 with 17 dog pictures gets mm07, D2 with 12 gets mm08, D3 with 8 gets mm09, and D4 with 3 gets mm10. The strongest qualitative result is that 3 dog pictures attract more traffic than 17 dog pictures under dynamic ranking. For ISP-oriented interpretation, this is best characterized as a direct algorithmic inverse-popularity effect rather than a vote–prediction ISP protocol.

6. Limits of the concept and open technical questions

The literature summarized here supports a broad ISP interpretation, but it also imposes clear boundary conditions. The ranking paper does not define an “inverse surprisingly popular” score distinct from surprisingly popular voting; it develops a practical ranking generalization via pairwise decomposition, ordinal prediction elicitation, and mm11-based conversion (Hosseini et al., 2021). The directed-network paper does not study respondents reporting both their own answer and their estimate of others’ answers, and it does not propose an answer-selection rule of the form “choose the option whose actual support exceeds predicted support”; its main mechanism is network exposure bias over a binary trait mm12 (Alipourfard et al., 2019). The popularity-ranking paper does not model Bayesian inference from meta-beliefs or truth discovery from minority-overexpected support; it studies sequential click behavior under popularity-based rankings (Germano et al., 2019).

A second limitation concerns theory. The asymptotic theorem quoted in the ranking paper applies to full SPV under the original model, not directly to the pairwise partial-information construction. The paper explicitly identifies open directions: guarantees for finite mm13, guarantees under partial votes and predictions, extension beyond four alternatives, and analysis under structured noise models such as the Mallows model. The popularity-ranking paper notes that for intermediate values

mm14

there may be multiple limit rankings, so the proof does not directly apply. The directed-network paper shows that local and global biases can differ in sign, which cautions against simplistic inversion procedures treating all “perceived popularity” measures as interchangeable.

A plausible synthesis is that ISP is best treated not as a single theorem but as a research program organized around three questions. The first asks how vote–prediction discrepancies can recover latent truths when informed minorities exist. The second asks how directed topology and attention allocation distort local prevalence estimates relative to global prevalence. The third asks how popularity-based visibility updates convert concentrated demand into disproportionate exposure. The papers considered here provide, respectively, a practical ranked-voting aggregation rule, a structural theory of neighborhood-based prevalence distortion with an explicit inversion strategy, and a rigorous algorithmic inverse-popularity effect. Together they show that inverse-popularity phenomena can arise from meta-beliefs, network structure, or ranking feedback, and that these mechanisms should not be conflated even when they produce similar empirical signatures.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Inverse Surprising Popularity (ISP).