Neighborhood Expectation Polling
- Neighborhood Expectation Polling is a network-sensitive method where respondents estimate the voting fraction in their social circle rather than their own vote.
- It uses neighborhood observations, network topology, and concepts like the friendship paradox to reduce estimator variance and improve inference accuracy.
- NEP is applied to predict electoral surprises, design adaptive polling in hierarchical networks, and correct data incest in social learning frameworks.
Searching arXiv for recent and foundational papers on Neighborhood Expectation Polling and closely related social circle polling. Neighborhood Expectation Polling (NEP) is a polling method on networks in which sampled individuals are asked not only for their own intention, but for an estimate of support in their neighborhood or social circle. In the canonical two-candidate formulation, respondents are asked for an estimate of the fraction of votes for a candidate; in the social-circle formulation, they are asked to estimate the voting intentions of their social contacts. Across the literature, NEP is analyzed as a local-to-global inference procedure whose statistical behavior depends on network structure, homophily, polarization, hierarchical influence, and information reuse. It has been studied as a mechanism for reducing mean squared error in network sampling, as a model of electoral surprise, as an adaptive polling primitive in hierarchical social networks, and as a setting in which misinformation propagation or “data incest” must be corrected (Nettasinghe et al., 2018, Dey et al., 2018, Palermo et al., 31 Mar 2025, Bhatt et al., 2018, Krishnamurthy et al., 2015).
1. Definition, question format, and relation to other polling schemes
NEP differs from both classical intent polling and expectation polling in the object it elicits. Classical intent polling asks sampled individuals for their own label or vote intention. Expectation polling asks who they think will win. NEP asks for an estimate of prevalence in a neighborhood: in one formulation, “What is your estimate of the fraction of people with label 1?”; in another, “Of all your social contacts who are likely to vote, what percentage will vote for [party/candidate]?” The response is therefore a neighborhood-level quantity rather than an individual intention or a binary winner forecast (Nettasinghe et al., 2018, Palermo et al., 31 Mar 2025).
This distinction matters because sampled individuals “naturally look at their neighbors” when answering the NEP question. In the network formulation of NEP, the response of individual is
and the aggregate estimator is
where is the sampled set. In the social-circle formulation, each respondent reports
the average voting intention among neighbors, and the corresponding estimator is
These formulations make explicit that NEP is a network-sensitive estimator rather than a simple opinion question (Nettasinghe et al., 2018, Palermo et al., 31 Mar 2025).
| Polling method | Typical question | Response object |
|---|---|---|
| Intent polling | “Who will you vote for?” | Own label or voting intention |
| Expectation polling | “Who do you think will win?” | Perceived winner |
| NEP / social circle polling | “What is your estimate of the fraction of votes for A?” / “What percentage will vote for [party/candidate]?” | Neighborhood or social-circle fraction |
A plausible implication is that NEP is best understood as a family of estimators whose performance is inseparable from the topology and information structure of the underlying network.
2. Local estimation of global preferences and electoral surprise
A central formalization of NEP appears in models of electoral surprise. In this line of work, individuals do not directly observe the global distribution of preferences; instead, they use their local neighborhood to infer it. The population is represented by a random graph , with connection probabilities biased by preferences. In the stochastic block model formulation, each voter belongs to one of classes corresponding to distinct preference orderings, and edge probabilities govern connections between classes. Within-class edges are more likely than between-class edges, so the model exhibits homophily (Dey et al., 2018).
Under this model, each voter estimates class sizes from observed neighbor counts and estimated connection probabilities. The paper defines surprise as a mismatch between the voter’s perceived winner and the true winner under a fixed voting rule: 0 For the two-candidate case, the decisive quantity is the discrepancy between the estimated bias and the true bias in local connections. Let 1 denote true connection probabilities, 2 a voter’s estimates, and let 3 encode the global vote-fraction skew. For a voter in class 1, surprise occurs with high probability as 4 if
5
and no surprise occurs with high probability if the inequality is reversed. Similar statements hold for class 2 voters. The paper also confirms the phenomenon that surprising outcomes are associated only with closely contested elections (Dey et al., 2018).
The multi-candidate extension compares plurality, Borda, and veto. Surprise depends both on a voter’s class and on the voting rule, and the analysis uses the most probable false beating (MPFB) factor. The results indicate that voting rules have different behavior for different parts of the population and hint at an impossibility that a single voting rule will be less surprising for all parts of a population. This suggests that, within NEP-style local inference, “debiasing” and rule choice are jointly consequential rather than separable design issues (Dey et al., 2018).
The empirical analysis based on the 2016 UK EU referendum uses a social graph with both class-based and geographic proximity. Each voter combines local observations and noisy global media information through a weighted average controlled by 6. The observed fraction of minority voters surprised varies with intra/inter class connection probabilities and the noise level in the global channel. High intra-class bias and high noise in global observation both increase the rate of surprise among minority voters, while more accurate estimates and higher weight on accurate global information decrease surprise (Dey et al., 2018).
3. Sampling design, friendship paradox, and mean squared error
A distinct research program treats NEP as an efficient polling method for estimating network-scale prevalence when the network is only partially observable. The key methodological ingredient is the friendship paradox: on average, a random friend has higher degree than a random node. Because higher-degree nodes are more central, sampling schemes that bias toward friends can reduce estimator variance. The paper studies two cases: the graph is unknown but random walks can be performed on it, and the graph is unknown apart from the ability to sample nodes uniformly (Nettasinghe et al., 2018).
Three algorithms are proposed. Algorithm 1 uses independent random walks and, after mixing, collects visited nodes as samples; Algorithm 2 uniformly samples nodes and then samples a random neighbor of each; the baseline “naive NEP” asks uniformly sampled nodes for the fraction of their own neighbors with label 1. Under iid Bernoulli labels, the paper gives the ordering
7
This is the sharpest formal statement in the paper of when friendship-paradox-based NEP dominates naive network polling (Nettasinghe et al., 2018).
For arbitrary labels, the analysis is explicitly bias–variance based. For random-walk NEP,
8
so bias is driven by degree-label correlation. The variance satisfies
9
where 0 is the second-largest singular value of the normalized adjacency matrix. For naive NEP, the variance is upper-bounded by
1
The paper emphasizes that degree-label correlation, network expansion, degree distribution, and assortativity all affect the mean squared error of NEP estimators (Nettasinghe et al., 2018).
Empirically, friendship-paradox-based NEP algorithms outperform classic intent polling and naive NEP in mean squared error for small sampling budgets, networks with little or no degree-label correlation, and heavy-tailed degree distributions. In Erdős–Rényi networks, by contrast, all methods are similar in performance because the friendship paradox is less pronounced. This suggests that NEP is not a single uniformly superior estimator; its efficiency gain is conditional on graph heterogeneity and label-degree dependence (Nettasinghe et al., 2018).
4. Network topology, polarization, and social-circle estimators
Recent work on “social circle polls” makes the connection between NEP and network topology explicit. In that framework, the population is an undirected graph 2, node labels are 3, and two structural parameters organize the analysis: average degree
4
and polarization
5
the fraction of edges that are cross-community. The election outcome is summarized by magnetization
6
Graphs are generated using the stochastic block model, enabling controlled variation in community size, polarization, and connectivity (Palermo et al., 31 Mar 2025).
The paper shows that the standard estimator 7 is unbiased, whereas the social circle estimator 8 is generally biased, with both bias and variance depending on topology, especially polarization 9. In simulations with 0, 1, and 2, the social circle estimator often has lower risk, measured by mean squared error, especially in less polarized, more connected networks and when the vote is close. In highly polarized networks, its performance degrades and can become as bad or worse than standard polling due to increased bias. Social circle polling is also more likely to predict the correct winner, especially in close elections with moderate network polarization (Palermo et al., 31 Mar 2025).
A major contribution is the construction of a topology-aware correction. Node heterophily is defined as
3
community heterophily averages are 4 and 5, and the estimated polarization is
6
Using this estimate, the adjusted estimator is
7
The paper states that this estimator is unbiased under the model and consistently outperforms both standard and uncorrected social circle polling for a wide parameter range, with degradation only in very sparsely connected networks 8 (Palermo et al., 31 Mar 2025).
The 2016 U.S. presidential election application uses GfK and USC Dornsife pre-election poll data with real respondent weights. In those single-survey cases, the adjusted estimator did not outperform the best of the basic estimators for both samples, possibly due to modeling assumptions, but it produced a polarization estimate of 9, interpreted as about 24% of acquaintances having contrary political preferences just before the election. A plausible implication is that topology estimation from poll responses can function both as a bias-correction device and as an internal validity check for NEP (Palermo et al., 31 Mar 2025).
5. Adaptive NEP in hierarchical social networks
In hierarchical social networks, NEP is generalized from a static estimator to an adaptive sensing mechanism. The paper “Adaptive Polling in Hierarchical Social Networks using Blackwell Dominance” refers to this extension as adaptive friendship polling. Nodes are organized in 0 levels, and influence flows from higher levels to lower levels. At each time, the pollster may select a level and ask sampled nodes for perceived fractional support for each state among their same-level neighbors. The underlying state of nature evolves as a finite-state Markov chain, so polling is directed toward tracking a time-varying latent state rather than estimating a fixed one-shot outcome (Bhatt et al., 2018).
The adaptive polling problem is formulated as a partially observed Markov decision process (POMDP). The unobserved state is 1, the action 2 selects a polling mechanism or level, the observation 3 is the polling outcome, and the belief state 4 is updated by Bayesian filtering: 5 The discounted cost objective is
6
For NEP, the observation likelihood is multinomial: when polling level 7, the probability of a reported neighbor-fraction vector is determined by the opinion distribution matrix 8 at that level (Bhatt et al., 2018).
The instantaneous cost is
9
where 0 is the measurement cost and the second term penalizes uncertainty in the state estimate. The central structural tool is Blackwell dominance. If 1 for all 2 and the cost is concave, then the myopic policy is an upper bound to the optimal policy: 3 When strict dominance is unavailable, the paper uses Le Cam deficiency to construct approximate Blackwell dominance and associated performance bounds. The same ordering also induces comparisons in mutual information and Rényi divergence for the resulting channels (Bhatt et al., 2018).
This formulation places NEP within sequential experimental design. It suggests that in hierarchical settings the main question is no longer only whether neighborhood reports are informative, but which level should be polled, at what cost, and with what effect on posterior uncertainty over time.
6. Data incest, social learning, and protocol design
A separate line of work analyzes expectation-style polling and NEP-like systems as instances of social learning on a directed acyclic graph (DAG). In this framework, nodes represent agents at particular times, edges represent influence, the adjacency matrix records one-hop influence, and the transitive closure matrix records reachability through multi-hop paths. The central pathology is “data incest”: unintentional re-use of identical actions in the formation of public belief, caused by overlapping information paths. In polling terms, the same informational source can be double-counted when respondents summarize their friends’ beliefs without accounting for common ancestry in the information flow (Krishnamurthy et al., 2015).
The goal of incest removal is to construct an incest-free posterior belief at node 4 using all available information exactly once. The paper gives a necessary and sufficient condition for exact incest removal using only immediate neighbors: 5 When the condition holds, the weights are
6
and the incest-free log-belief is formed as
7
followed by exponentiation, normalization, and Bayesian fusion with private observation. If the condition fails, fair estimation is not possible using only immediate friends; additional information must be collected, or one must compute the best possible unbiased estimator from incestious beliefs, typically at higher variance (Krishnamurthy et al., 2015).
The experimental evidence in the paper is unusually direct. In a human-subject study with 36 undergraduate students and 1658 trials involving dyads, “data incest” patterns occurred in 79% of trials, and in 21% of trials those patterns accounted for decision changes. Herding occurred in 66% of trials, but only 32% converged correctly. The paper also uses revealed preference theory and Twitter datasets to argue that social sensors are utility maximizers, thereby supporting the behavioral assumptions underlying social-learning-based polling models (Krishnamurthy et al., 2015).
For NEP, the significance is methodological rather than merely cautionary. It suggests that network-aware polling requires not only debiasing for homophily or polarization, but also explicit control of information reuse. In settings where reports already incorporate others’ reports, estimator design and communication protocol design become inseparable.