Inverse Surprising Popularity
- Inverse Surprising Popularity (ISP) is a phenomenon where realized, predicted, and perceived support diverge, leading to a distorted signal of true prevalence.
- The ranked-voting approach uses pairwise reduction with partial votes and ordinal predictions to recover true rankings despite noisy, inverse-popular signals.
- Additional models show that network exposure bias and popularity-based ranking feedback can concentrate support, exemplified by the few-get-richer effect in social media settings.
Inverse Surprising Popularity (ISP) is not defined by name in the papers considered here. In this literature, the term is most naturally understood as referring to settings in which realized support, predicted support, and locally perceived support diverge systematically, so that naive popularity is a distorted signal of a latent global truth, ranking, or prevalence. That interpretation is explicit in surprisingly popular voting and its ranking-recovery extension, where respondents report both their own answer and a prediction of others’ answers (Hosseini et al., 2021). It is closely mirrored, though not formally identical, in directed-network models of prevalence distortion induced by the friendship paradox (Alipourfard et al., 2019) and in popularity-based ranking systems exhibiting the few-get-richer effect (Germano et al., 2019). Taken together, these works place ISP at the intersection of meta-belief aggregation, network exposure bias, and algorithmic visibility dynamics.
1. Conceptual scope and relation to surprisingly popular voting
The most direct formal lineage for ISP is classical surprisingly popular voting. In that setting, each participant reports a vote and a prediction of other participants’ votes, and the aggregation rule selects the option whose actual support exceeds its predicted support in the relevant sense. The ranking paper restates the classical motivation succinctly: classical democratic aggregation works when the majority is relatively accurate, whereas surprisingly popular voting can recover the ground truth even when experts are in minority (Hosseini et al., 2021).
The three papers considered here do not describe a single unified ISP formalism. One paper develops a ranking-recovery analogue of classical surprisingly popular voting; one analyzes neighborhood-based prevalence distortion in directed networks; and one studies an algorithmic inverse-popularity effect under popularity-based ranking. A common source of confusion is to treat these as the same object. The literature instead separates at least three distinct mechanisms. First, there is the canonical vote–prediction discrepancy mechanism of surprisingly popular voting. Second, there is directed-network perception bias, where local exposure overweights structurally prominent actors. Third, there is ranking-feedback amplification, where popularity-updated visibility can cause a smaller class of items to attract more aggregate traffic than a larger class.
This distinction matters for ISP because the phrase “inverse popularity” can refer either to an aggregation rule that corrects naive vote counts by using meta-predictions, or to a structural process in which visibility and exposure make observed popularity systematically misleading. The ranking paper is the closest to ISP in protocol design; the directed-network and popularity-ranking papers are best understood as mechanistic analogies or network-structural foundations rather than direct implementations of the standard surprisingly popular method (Alipourfard et al., 2019, Germano et al., 2019).
2. Pairwise surprisingly popular aggregation for ranking recovery
The most explicit extension of surprisingly popular logic beyond single-option truth discovery is the ranked-voting framework of “Surprisingly Popular Voting Recovers Rankings, Surprisingly!” (Hosseini et al., 2021). Let be a set of alternatives, the set of rankings over , and the true ranking. Each voter observes a noisy ranking , drawn from a signal distribution conditional on . The model uses the posterior
and from that derives another voter’s signal distribution,
If full votes and full predictive distributions over rankings were elicited, the original SP score is
0
Under the condition
1
the paper quotes the theorem from Prelec et al.: 2
The practical difficulty is that direct SPV over rankings requires predictions over all 3 rankings. The paper therefore studies partial votes and partial predictions. The vote formats are Top vote and Rank vote. The prediction formats are Top prediction and Rank prediction. This yields the four principal elicitation formats Top-Top, Top-Rank, Rank-Top, and Rank-Rank, together with the no-prediction baselines Top-None and Rank-None.
The methodological core is a pairwise reduction. For each pair 4, the method extracts a binary pairwise vote 5 and a pairwise prediction 6. Because classical SPV requires cardinal predictions but the elicited predictions are ordinal, the paper introduces a two-parameter conversion using
7
These parameters are learned from training data. The pairwise SP-style scores are then
8
and
9
If 0, the method records 1; otherwise it records 2.
Because pairwise decisions are made independently, the output can be a cyclic tournament rather than a globally coherent ranking. The paper therefore uses the resulting tournament in two ways: predicting the top alternative by choosing the alternative that defeats the largest number of others, and evaluating full-ranking recovery by Kendall Tau distance to the ground truth. The empirical study uses 720 MTurk participants and 7,200 responses over geography, movies, and paintings tasks with objective ground-truth rankings. The central findings are that prediction helps substantially, Rank-Rank performs best, and even a little prediction information helps surprisingly popular voting outperform classical approaches. Particularly notable is the result that Top-Rank significantly outperforms Rank-Top, suggesting that richer prediction elicitation can matter more than richer vote elicitation.
3. Directed-network prevalence distortion as an ISP-adjacent mechanism
“Friendship Paradox Biases Perceptions in Directed Networks” formalizes a different but closely related object: the gap between true global prevalence and locally perceived prevalence in a directed network (Alipourfard et al., 2019). The network is 3, with 4. A directed edge 5 means that 6 is a friend of 7, or equivalently that 8 follows 9; the edge direction is the direction of information flow. For a node 0, 1 denotes out-degree, interpreted as number of followers, and 2 denotes in-degree, interpreted as number of friends.
The paper defines three sampling-based random variables. A random node 3 is uniform over nodes: 4 A random friend 5 is sampled proportional to out-degree: 6 A random follower 7 is sampled proportional to in-degree: 8 Because total in-degree equals total out-degree,
9
Each node has a binary attribute 0. The true global prevalence is 1. The global feed-level perceived prevalence is 2. For node 3, the local perceived prevalence is
4
and the average local perception is 5. This decomposition is especially important for ISP-style reasoning because it distinguishes actual prevalence, friend-weighted prevalence, and neighborhood-level prevalence.
A major result is the explicit bias formula
6
The paper also rewrites this bias in terms of the Pearson correlation 7, out-degree standard deviation 8, and binary-attribute standard deviation 9. The interpretation is direct: if trait-holders have higher out-degree, they are overrepresented in others’ feeds, and degree heterogeneity strengthens the distortion.
The paper further defines local perception bias,
0
and introduces attention
1
It derives
2
Sufficient conditions for positive local perception bias are
3
and
4
The appendix states that these conditions imply
5
The same paper formalizes four directed variants of the friendship paradox. Two hold in any directed network: random friends have more followers than random nodes, and random followers have more friends than random nodes. Two additional variants require positive in/out-degree correlation: random friends have more friends than random nodes, and random followers have more followers than random nodes. A plausible implication for ISP is that “perceived popularity” is not a single object. The paper explicitly emphasizes that global and local biases can even have opposite signs.
4. Estimation from biased perceptions in directed networks
The same directed-network framework also provides a concrete estimation-from-biased-perceptions procedure through Follower Perception Polling (FPP) (Alipourfard et al., 2019). The goal is to estimate
6
the true global prevalence. FPP samples 7 independent nodes from the follower distribution
8
and uses the estimator
9
The motivation is a bias–variance tradeoff. Random followers tend to have larger in-degree than random nodes, so they observe more friends, and their neighborhood averages 0 have lower variance than the perceptions of random nodes. The paper’s bias theorem states
1
From the appendix,
2
so FPP intentionally accepts the same bias as friend-weighted prevalence rather than directly debiasing local perceptions.
The paper also gives an explicit unbiased modification,
3
Methodologically, this is the clearest inversion step in the directed-network setting: degree-weighted local information can be corrected by inverse-degree weighting.
For variance, let 4 be the adjacency matrix, 5, and 6. The degree-discounted bibliographic coupling matrix is
7
If 8 is connected and non-bipartite, the paper provides an exact variance expression and the upper bound
9
where 0 and 1 is the second largest eigenvalue of 2.
The empirical validation uses Twitter data from 2014. Starting from 100 users active on California ballot initiatives in 2012, the collection expands to 5,599 seed users, then to over 600K users total, with over 18M hashtag mentions. Each hashtag 3 is treated as a binary trait 4 if user 5 used hashtag 6. Among the 1,153 most popular hashtags, each used by more than 1,000 people, 865 had positive local bias. For #ferguson, global prevalence is 7 and local perceived prevalence is 8, so it appeared about four times more popular than it actually was. In synthetic polling experiments on an induced subgraph of 5,409 users, FPP has lower variance than both Intent Polling and Node Perception Polling, and for sample budget 9 it outperforms both in MSE for most hashtags. Even at 0, it outperforms IP in more than 80% of cases and NPP in more than 55% of cases.
5. Popularity-based ranking and the few-get-richer effect
“The few-get-richer: a surprising consequence of popularity-based rankings” analyzes a distinct inverse-popularity mechanism in search and recommender systems (Germano et al., 2019). There are 1 items partitioned into two classes, 2 and 3. Sequentially arriving users have heterogeneous class preferences. For user 4, the item-level propensity absent ranking is
5
This assumption is central: class-level preference is divided uniformly across items within a class, so fewer items imply more concentrated per-item support.
Popularity-based ranking updates according to cumulative clicks. With attention bias parameter 6, the click probability is
7
Aggregate class-level traffic share is
8
The paper’s central result is the few-get-richer effect: the fewer the items in a given class, the higher the total share of traffic that class may collectively attract. The analytical thresholds are sharp. If
9
a stable limit ranking places all class-1 items at the top. If
0
class 1 is sufficiently large that class 0 is the small class, and the 1 items are bottom-ranked. The formal proposition compares two environments 2 and 3 differing only in 4 and 5, and states that there exists 6 such that for any 7, total clicking probability on class 1 is strictly greater in the environment with fewer class-1 items, provided 8 is sufficiently small: 9
Mechanistically, concentrated niche demand lifts a small number of minority-class items in the ranking; once they reach high positions, indifferent users and rank-biased users reinforce them. The paper’s simulations examine 00, 01, uniform initialization, and varying 02, 03, and 04. The reported pattern is that CTR to 05 decreases as 06 increases, and stronger attention bias strengthens the effect.
The online experiment uses 786 Amazon Mechanical Turk participants. Participants identify as cat person, dog person, or neither, then choose among 20 clickable cat or dog pictures. There are four dynamic conditions with popularity-updated ranking and four static matched controls. In the dynamic conditions, dog pictures receive more than 50% of traffic in all cases: D1 with 17 dog pictures gets 07, D2 with 12 gets 08, D3 with 8 gets 09, and D4 with 3 gets 10. The strongest qualitative result is that 3 dog pictures attract more traffic than 17 dog pictures under dynamic ranking. For ISP-oriented interpretation, this is best characterized as a direct algorithmic inverse-popularity effect rather than a vote–prediction ISP protocol.
6. Limits of the concept and open technical questions
The literature summarized here supports a broad ISP interpretation, but it also imposes clear boundary conditions. The ranking paper does not define an “inverse surprisingly popular” score distinct from surprisingly popular voting; it develops a practical ranking generalization via pairwise decomposition, ordinal prediction elicitation, and 11-based conversion (Hosseini et al., 2021). The directed-network paper does not study respondents reporting both their own answer and their estimate of others’ answers, and it does not propose an answer-selection rule of the form “choose the option whose actual support exceeds predicted support”; its main mechanism is network exposure bias over a binary trait 12 (Alipourfard et al., 2019). The popularity-ranking paper does not model Bayesian inference from meta-beliefs or truth discovery from minority-overexpected support; it studies sequential click behavior under popularity-based rankings (Germano et al., 2019).
A second limitation concerns theory. The asymptotic theorem quoted in the ranking paper applies to full SPV under the original model, not directly to the pairwise partial-information construction. The paper explicitly identifies open directions: guarantees for finite 13, guarantees under partial votes and predictions, extension beyond four alternatives, and analysis under structured noise models such as the Mallows model. The popularity-ranking paper notes that for intermediate values
14
there may be multiple limit rankings, so the proof does not directly apply. The directed-network paper shows that local and global biases can differ in sign, which cautions against simplistic inversion procedures treating all “perceived popularity” measures as interchangeable.
A plausible synthesis is that ISP is best treated not as a single theorem but as a research program organized around three questions. The first asks how vote–prediction discrepancies can recover latent truths when informed minorities exist. The second asks how directed topology and attention allocation distort local prevalence estimates relative to global prevalence. The third asks how popularity-based visibility updates convert concentrated demand into disproportionate exposure. The papers considered here provide, respectively, a practical ranked-voting aggregation rule, a structural theory of neighborhood-based prevalence distortion with an explicit inversion strategy, and a rigorous algorithmic inverse-popularity effect. Together they show that inverse-popularity phenomena can arise from meta-beliefs, network structure, or ranking feedback, and that these mechanisms should not be conflated even when they produce similar empirical signatures.