Papers
Topics
Authors
Recent
Search
2000 character limit reached

Collective Confidence in Group Decisions

Updated 15 July 2026
  • Collective confidence is the group-level measure of uncertainty and reliability derived from aggregating individual judgments, applicable in diverse domains such as crowd inference and macroeconomics.
  • Normative aggregation methods like confidence-weighted majority voting and log-odds pooling outperform simple averages, with simulations showing improvements (e.g., 76.2% vs 66.7% accuracy).
  • Recent studies in social influence and multi-agent LLM systems reveal that calibrated collective confidence enhances consensus, mitigates overconfidence, and improves factual accuracy.

Searching arXiv for recent and foundational papers on collective confidence, confidence-weighted aggregation, social influence, and multi-agent deliberation. Collective confidence is the group-level counterpart of confidence: a representation of uncertainty, reliability, or plausibility attached to a shared judgment, estimate, proposition, partition, or macro-level expectation. In current research, it is not a single construct but a family of formalisms. In small-group decision studies, it is the probability that a group decision is correct; in crowd inference, it is the pooled probability that a proposition is true; in social influence research, it is a cue that modulates who influences whom; in uncertainty-aware annotation, it is the distributional structure of collective human judgments; in panel econometrics, it is a confidence set over joint group assignments; and in multi-agent LLM systems, it is a reliability signal emerging from calibrated confidence, diversity, agreement, and interaction dynamics (Meyen et al., 2020, Stinson et al., 9 Jan 2025, Wang et al., 2023, Dzemski et al., 2017, Zhu et al., 9 Jan 2026).

1. Core meanings and operationalizations

A central distinction in the literature is between individual confidence and collective confidence. In group decision experiments, an individual makes a binary decision yi{+1,1}y_i\in\{+1,-1\} and reports a confidence

ci=P(yi is correct),c_i = P(y_i \text{ is correct}),

so confidence is treated as a subjective probability of correctness rather than a verbal label or ordinal rank. A real group then reports a group confidence cgc_g, interpreted analogously as the probability that the group decision is correct (Meyen et al., 2020).

In crowd inference on propositions, the primitive object is not a binary vote but a continuous probability judgment on [0,1][0,1]. The collective problem is to infer proposition truth from a matrix of peer ratings while accounting for differences in peer accuracy and calibration. Here collective confidence is a pooled probability attached to a claim, and the literature emphasizes that continuous probability judgments contain more information than thresholded yes/no votes (Stinson et al., 9 Jan 2025).

In social influence research, confidence is not merely an output variable. It is a signal that changes influence asymmetrically: individuals revise more when the other person is equally or more confident, and resist feedback when already more confident. The resulting collective confidence can track consensus rather than truth once interaction begins, because corroboration raises confidence even when the corroborating opinion is inaccurate (Moussaid et al., 2013).

Other fields generalize the notion further. In semantic textual similarity, “collective confidence” refers to how tightly or loosely annotators agree on a sentence pair, so the collective object is a full opinion distribution rather than a single score (Wang et al., 2023). In grouped panel models, it refers to a confidence set over the full latent assignment vector (g1,,gN)(g_1,\dots,g_N), with joint coverage over all units rather than a probability of correctness for one decision (Dzemski et al., 2017). In macroeconomics, consumer confidence is treated as an emergent collective psychological state that can be inferred from large-scale online search behavior (Dong et al., 2014).

These usages share a common theme: confidence becomes collective when uncertainty is attached to an aggregate object and when aggregation, interaction, or inference rules determine how individual uncertainty enters the group-level quantity.

2. Normative aggregation of judgments and probabilities

The most explicit normative model is confidence-weighted majority voting. In the small-group decision framework of binary judgments, unweighted majority voting ignores confidence and uses

ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).

Confidence-weighted majority voting instead defines

wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),

and predicts collective confidence by

cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.

Under independence of individual decisions and accurate confidence estimates, this is the mathematically optimal aggregation rule. Empirically, simulated group decisions based on CWMV matched real group accuracy very closely: real groups were 76.2%76.2\% correct, CWMV simulations were 76.2%76.2\% correct, and MV simulations were ci=P(yi is correct),c_i = P(y_i \text{ is correct}),0 correct. Real groups nevertheless showed an equality bias and underconfidence relative to the normative benchmark, with fitted parameters ci=P(yi is correct),c_i = P(y_i \text{ is correct}),1 and ci=P(yi is correct),c_i = P(y_i \text{ is correct}),2 (Meyen et al., 2020).

An analogous log-odds logic appears in proposition-level crowd inference. If ratings are well-calibrated and conditionally independent, the independent opinion pool yields

ci=P(yi is correct),c_i = P(y_i \text{ is correct}),3

This uses the confidence structure of graded judgments rather than discarding it. The same work shows that individually calibrated aggregation outperforms uncalibrated rating averaging, group-level calibration, and majority vote, and that an unsupervised EM-based generative model can infer peer accuracy and miscalibration jointly with claim probabilities, reaching about ci=P(yi is correct),c_i = P(y_i \text{ is correct}),4 accuracy with enough ratings per claim (Stinson et al., 9 Jan 2025).

The literature also studies aggregation rules in which confidence affects who participates. In a confidence-gated generalization of the Condorcet Jury Theorem, agents update a Beta posterior over their own competence during a calibration phase and vote only if their posterior confidence exceeds an abstention threshold. The collective success probability is then lower-bounded in terms of the expected contributions of the publishing agents, and asymptotic reliability is recovered provided competent agents are not filtered out with vanishing probability (Karge, 25 Feb 2026). This suggests that collective confidence may improve not only by weighting votes, but also by allowing calibrated abstention.

A related strategic result appears in French-DeGroot dynamics with endogenous self-confidence. There, each agent chooses a self-confidence level ci=P(yi is correct),c_i = P(y_i \text{ is correct}),5, and the Pareto-optimal profiles satisfy

ci=P(yi is correct),c_i = P(y_i \text{ is correct}),6

Under these profiles, the emergent consensus weights coincide with the minimum-variance unbiased linear estimator, and the same profiles are strict Nash equilibria in the interior of the strategy space (Como et al., 22 Jun 2025). Collective confidence is thus linked to precision, social power, and self-weighting simultaneously.

3. Confidence as a driver of social dynamics

Research on social influence shows that collective confidence is not a passive summary of opinions. It is a dynamical variable that changes who moves, how much they move, and whether a group converges. In controlled experiments, subjects first answered factual questions and then revised their responses after observing a peer’s estimate and confidence. Across 885 dyadic interactions, revisions fell into three heuristics—keep own opinion, compromise, or adopt the other opinion—with overall frequencies of ci=P(yi is correct),c_i = P(y_i \text{ is correct}),7, ci=P(yi is correct),c_i = P(y_i \text{ is correct}),8, and ci=P(yi is correct),c_i = P(y_i \text{ is correct}),9, respectively. The influence map is organized by confidence difference cgc_g0 and normalized opinion distance. It exhibits a confirmation zone for cgc_g1, an influence zone for cgc_g2, and a far zone for cgc_g3. The same work identifies two attractors of collective opinion: the expert effect, induced by a highly confident individual, and the majority effect, induced by a critical mass of laypeople sharing similar opinions. Simulations show a tipping point around cgc_g4 experts; below roughly cgc_g5, the majority effect dominates, whereas in the cgc_g6–cgc_g7 range the collective outcome shifts toward expert opinion (Moussaid et al., 2013).

This line of work also supports a more cautionary interpretation: confidence initially tracks accuracy, but under social interaction it increasingly tracks consensus rather than truth. Collective confidence can therefore be socially amplified independently of objective correctness (Moussaid et al., 2013).

A complementary result appears in the “wisdom of the confident” framework. Here confidence is inferred behaviorally rather than self-reported. Individuals first provide an estimate cgc_g8, then revise it to cgc_g9 after receiving social information [0,1][0,1]0. The inferred social weight

[0,1][0,1]1

serves as a proxy for confidence in private information: low [0,1][0,1]2 indicates relative resistance to social influence. Aggregating only the first estimates of low-[0,1][0,1]3 individuals with the geometric mean yields an unbiased trend, in contrast to using the full-group median. For the Switzerland–Italy border length question, the true value was [0,1][0,1]4, whereas the full-group wisdom-of-the-crowd estimate was [0,1][0,1]5 (Polavieja et al., 2014). The implication is not that social interaction is uniformly harmful, but that interaction reveals heterogeneity that can itself be exploited for aggregation.

Bounded-confidence models add a topological dimension. In the Hegselmann–Krause model, agents average only over neighbors whose opinions are within a confidence radius [0,1][0,1]6. On sparse, unbiased networks, collective behavior depends jointly on [0,1][0,1]7 and the normalized edge density [0,1][0,1]8. The phase diagrams show fragmentation at low [0,1][0,1]9, larger consensus regions at high (g1,,gN)(g_1,\dots,g_N)0, and a nonmonotonic re-entrant transition in which increased connectivity can suppress consensus rather than promote it. For large systems, the critical confidence threshold stabilizes near (g1,,gN)(g_1,\dots,g_N)1, but low connectivity prevents full unanimity and significantly alters the phase boundaries in finite populations (Molignini, 18 Jul 2025). This challenges the misconception that more communication automatically strengthens collective agreement.

4. Collective confidence in multi-agent LLM systems

Recent work on LLM collectives treats confidence as an explicit control variable for answer aggregation, deliberation, and error suppression. A training-free method called Collaborative Calibration simulates a two-stage group deliberation process. In the first stage, specialized agents using Chain-of-Thought, Program-of-Thoughts, Search-Augmented Self-Ask, and GenRead generate answers and initial confidence estimates; semantically equivalent answers are then clustered into stances. In the second stage, more general agents defend stances, rate arguments on consistency, clarity, conciseness, and factuality, revise their answers, and re-estimate confidence with a rationale. On six generative QA tasks—GSM8K, TriviaQA, SciQ, AmbigQA, DateUnd, and Biz-Ethics—the method achieves lower ECE on 4 out of 6 tasks, especially GSM8K, DateUnd, and AmbigQA, while showing similar or better Brier scores without hurting accuracy or generation quality (Yang et al., 2024).

In multi-agent debate, a sharper theoretical distinction has been drawn between diversity and confidence. Under homogeneous, unweighted debate dynamics, the belief in the correct option is a martingale, so debate does not systematically improve the expected probability of correctness. Introducing explicit confidence changes this. In a Dirichlet–categorical model, each output becomes (g1,,gN)(g_1,\dots,g_N)2, with (g1,,gN)(g_1,\dots,g_N)3, and raw answer counts are replaced by confidence-weighted counts. If confidence is positively correlated with correctness, the resulting process becomes a strict submartingale: (g1,,gN)(g_1,\dots,g_N)4 The same work reports substantial calibration gains after confidence training: for Llama-3.1-8B-Instruct, the Brier score drops from (g1,,gN)(g_1,\dots,g_N)5 to (g1,,gN)(g_1,\dots,g_N)6 and ECE from (g1,,gN)(g_1,\dots,g_N)7 to (g1,,gN)(g_1,\dots,g_N)8; for Qwen-2.5-7B-Instruct, Brier drops from (g1,,gN)(g_1,\dots,g_N)9 to ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).0 and ECE from ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).1 to ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).2. Across six reasoning-oriented QA benchmarks, diversity-aware initialization plus confidence consistently outperforms vanilla debate and majority vote (Zhu et al., 9 Jan 2026).

Other work addresses settings without ground truth. A multi-LLM collective reasoning pipeline rotates question-generation and answering roles among GPT-4-0125-preview, LLaMA-3-70B-Instruct, Claude-3-Opus, and Gemini-1.5-Flash. In that framework, agreement among independent models serves as a surrogate for confidence and question quality. Collective confidence is quantified with chi-square tests, Fleiss’ Kappa, and bootstrap confidence intervals. Gemini reaches ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).3, Claude reaches ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).4, GPT-4 reaches ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).5, and LLaMA reaches ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).6. The reported bootstrap confidence intervals for consensus rates are ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).7–ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).8 for Claude, ygMV=sign(i=1nyi).y_g^{MV} = \text{sign}\left(\sum_{i=1}^n y_i\right).9–wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),0 for GPT-4, wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),1–wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),2 for Gemini, and wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),3–wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),4 for LLaMA (Davoudi et al., 28 Feb 2025). In this setting, confidence is collective and relational rather than an internal uncertainty score of a single model.

Confidence becomes even more consequential in recursive multi-agent systems subject to hallucination. In a network model where outputs are decomposed into claims wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),5 with confidence scores wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),6, the confidence-weighted hallucination rate penalizes unsupported claims more heavily when they are stated confidently. Propagation is governed by confidence-coupled adoption probabilities and a time-varying impact matrix. The proposed defense combines confidence-weighted aggregation, adaptive impact regulation, external verification, and selective isolation of unreliable agents; final consensus is formed by a trust-weighted vote. On TruthfulQA and TriviaQA, the method reduces hallucination by up to wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),7, improves factual accuracy from wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),8 to wi=log ⁣(ci1ci),ygCWMV=sign(i=1nwiyi),w_i = \log\!\left(\frac{c_i}{1-c_i}\right), \qquad y_g^{\text{CWMV}} = \text{sign}\left(\sum_{i=1}^n w_i y_i\right),9, improves semantic consistency from cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.0 to cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.1, and limits hallucination amplification to cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.2 under adversarial conditions, compared with cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.3 without adaptive control (Jamshidi, 6 Jun 2026).

5. Collective confidence as disagreement and statistical uncertainty

In annotation-heavy tasks, collective confidence is often better understood as the shape of human disagreement than as a single average confidence score. The USTS dataset for semantic textual similarity contains about cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.4 Chinese sentence pairs and more than cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.5 labels. It separates uncontroversial cases, USTS-U, from contentious cases, USTS-C, using the threshold cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.6. USTS-U contains cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.7 examples with mean cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.8 and annotator correlation cgCWMV=11+exp ⁣(i=1nwiyi).c_g^{\text{CWMV}}=\frac{1}{1+\exp\!\left(-\left|\sum_{i=1}^n w_i y_i\right|\right)}.9; USTS-C contains 76.2%76.2\%0 examples with mean 76.2%76.2\%1 and annotator correlation 76.2%76.2\%2 (Wang et al., 2023).

The central finding is that averaging masks the true distribution of human opinions on low-agreement examples. A single scalar mean cannot distinguish tightly clustered judgments from bimodal or otherwise dispersed judgments with the same average. A single Gaussian is also often inadequate: on test data, a multimodal GMM outperforms a unimodal Gaussian on 76.2%76.2\%3 of instances. At the same time, the evidence for strong multimodality is limited—only about 76.2%76.2\%4 of pairs have a meaningful second component with weight at least 76.2%76.2\%5, and only 3 pairs have a meaningful third component (Wang et al., 2023). Current STS models do not capture this instance-level collective uncertainty well: model-predicted variance shows essentially no meaningful correlation with human disagreement, indicating that model uncertainty reflects predictive confidence over the aggregate dataset rather than collective human confidence on specific items (Wang et al., 2023).

A different statistical tradition uses confidence in the classical inferential sense. In grouped panel models, the object of interest is a latent assignment vector 76.2%76.2\%6. The confidence set 76.2%76.2\%7 is required to satisfy the asymptotic joint coverage property

76.2%76.2\%8

The construction proceeds by inverting many simultaneous unit-specific one-sided tests for group membership, forming marginal unit sets 76.2%76.2\%9, and taking the Cartesian product 76.2%76.2\%0. Bonferroni correction yields joint coverage, and Monte Carlo evidence indicates coverage near or above the nominal 76.2%76.2\%1 in finite samples (Dzemski et al., 2017). Here collective confidence is neither consensus nor correctness probability; it is formal uncertainty quantification for an entire estimated partition.

6. Macroscopic indicators, recurring misconceptions, and broader implications

At the macro level, collective confidence can denote a population-wide psychological state inferred from behavioral traces. The Chinese Consumer Confidence Index (C3I) uses monthly official survey data from January 2006 to June 2013 together with Google Trends query volumes derived from the Economist’s Confidence Questionnaire. Of 44 initial topics, 34 had sufficient non-zero search volume and were retained. Principal component analysis on these 34 series kept the first 9 components, explaining about 76.2%76.2\%2 of the variance. The fitted regression for C3I achieved 76.2%76.2\%3 and adjusted 76.2%76.2\%4, passed diagnostic tests for heteroskedasticity and cointegration, and tracked out-of-sample data from July 2013 to May 2014 well despite Google renormalization (Dong et al., 2014). In this usage, confidence is not elicited from respondents directly but inferred from collective online attention.

Several misconceptions recur across these literatures. The first is that collective confidence is equivalent to a simple average or a simple majority. The evidence is repeatedly against this: CWMV outperforms unweighted majority voting when confidence is informative (Meyen et al., 2020), continuous logit-based aggregation outperforms binary voting when probability judgments are available (Stinson et al., 9 Jan 2025), and average STS labels can erase the actual structure of human disagreement (Wang et al., 2023).

The second is that confidence necessarily tracks truth. Social influence experiments show that confidence can become a signal of consensus rather than accuracy (Moussaid et al., 2013); bounded-confidence dynamics show that consensus itself depends on network structure and can be destabilized by increased connectivity (Molignini, 18 Jul 2025); and multi-agent LLM systems show that high-confidence errors can propagate, reinforce, and amplify unless confidence is calibrated and interaction is regulated (Jamshidi, 6 Jun 2026).

The third is that consensus alone is sufficient. Agreement among models can function as a practical reliability proxy when no ground truth exists, but this is explicitly a surrogate measure rather than correctness in the classical sense (Davoudi et al., 28 Feb 2025). Conversely, calibrated confidence communication can improve aggregation when it is tied to diversity, verification, or abstention rules (Zhu et al., 9 Jan 2026, Yang et al., 2024, Karge, 25 Feb 2026).

Taken together, the literature presents collective confidence as a unifying but nonuniform concept. It may denote a normative aggregation rule, a social influence signal, a distribution of human judgments, a frequentist uncertainty region, an emergent macro-psychological index, or an interaction-level control variable in AI systems. The common analytical question is always the same: how should uncertainty expressed by multiple agents be represented, propagated, weighted, or filtered so that the collective output is not merely decisive, but appropriately calibrated to what the group actually knows.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Collective Confidence.