---
title: Subset-Based Collective Decision-Making
url: https://www.emergentmind.com/topics/subset-based-collective-decision-making-subcdm
type: topic
---

# Subset-Based Collective Decision-Making

to=arxiv_search  天天中彩票双色球json
{"query":"2307.06077 OR 2508.00467 OR 1204.3860 OR 2112.00574 OR 2207.04983 OR 1612.00071 OR 2103.14351 OR 2509.05182 OR 1410.5738 OR 2308.03376","max_results":10,"sort_by":"submittedDate"} code
to=arxiv_search ՞նչjson
{"query":"2307.06077","max_results":3,"sort_by":"relevance"} code
to=arxiv_search  үрүмjson
{"query":"A Generalised Theory of Proportionality in Collective Decision Making","max_results":5} code
to=arxiv_search 】【。】【”】【json
{"search_query":"A Generalised Theory of Proportionality in Collective Decision Making","max_results":5} code
Subset-Based Collective Decision-Making (SubCDM) designates a class of collective-choice models in which either the outcome is a subset chosen under constraints or the decision process itself is restricted to subsets of agents, alternatives, issues, or observations. In the literature, the term does not denote a single formalism. Rather, it appears in several technically distinct settings: proportional representation over feasible subsets of candidates, decentralized swarm consensus using only a subset of robots, collective computation under subset-based information allotments, stochastic dynamics driven by sampled neighborhoods or pairwise subsets, and learning or optimization over subset-valued alternatives [2307.06077] [2508.00467] [1204.3860].

## 1. Conceptual scope and recurring structure

Across these strands, SubCDM replaces unrestricted full-participation aggregation with structured partiality. The subset may be the chosen outcome, the active decision-making population, the locally visible data, or the interaction neighborhood used in an update rule.

| Strand | Subset object | Canonical formalism |
|---|---|---|
| Social choice | Feasible outcome subset \(W \subseteq C\) | \(F \subseteq 2^C\), BEJR/EJR/PJR |
| Swarm robotics | Decision-making robot subset \(S_{DM}\) | Hop-based or probabilistic recruitment |
| Communication complexity | Allotment subsets \(S_i \subseteq [n]\) | Macroscope \((f,A)\) |
| Nonlinear and stochastic dynamics | Sampled pairs, neighborhoods, or hyperedges | Urn process, hypergeometric sampling, hypernetwork ODEs |
| Preference learning and optimization | Subsets as alternatives or feasible judgments | Robust ordinal regression; CDO as judgment aggregation |

A recurring misconception is that subset-based decision-making is simply a lossy approximation to full participation. The surveyed literature presents a more differentiated picture. In social choice, the subset is the outcome and proportionality is strengthened rather than weakened. In swarm robotics, restricting active participation is explicitly intended to preserve accuracy while reducing resource use. In communication-complexity models, subset structure and meta-information can reduce communication from input-scale disclosure to task-specific summaries. In nonlinear dynamics, higher-order subset interactions can change the bifurcation structure of the decision process itself.

## 2. Social-choice foundations: proportionality over cohesive voter subsets

The most systematic formalization of subset-based collective decision-making in social choice models a collective outcome as a feasible subset of items. Voters are \(N=\{1,\dots,n\}\), items are \(C=\{c_1,\dots,c_m\}\), and \(F \subseteq 2^C\) is a nonempty family of feasible sets, assumed closed under inclusion. Each voter \(i\) has approval set \(A_i \subseteq C\), with utility \(u_i(W)=|W \cap A_i|\). This framework unifies committee elections, public decisions, diversity-constrained selection, and collective scheduling [2307.06077].

The central normative move is to define proportionality directly for arbitrary subsets \(S \subseteq N\), without predefining demographic groups. A subset \(S\) deserves \(\ell\) under Base Extended Justified Representation (BEJR) if, for every \(T \in F\), either there exists \(X \subseteq \bigcap_{i \in S} A_i\) with \(|X| \ge \ell\) such that \(T \cup X \in F\), or
\[
\frac{|S|}{n} > \frac{\ell}{|T|+\ell}.
\]
An outcome \(W \in F\) satisfies BEJR if every subset deserving \(\ell\) contains some voter \(i\) with \(u_i(W) \ge \ell\). Extended Justified Representation (EJR) strengthens this by conditioning the claim on the actual outcome \(W\): for every \(T \subseteq W\), the same feasibility-or-proportional-overruling condition must hold. The framework also defines PJR and BPJR generalizations, with PJR using \(\ell = |(\bigcup_{i\in S}A_i)\cap W|+1\).

These axioms admit constructive rule adaptations. Proportional Approval Voting is generalized by maximizing
\[
\mathrm{PAV}(W)=\sum_{i\in N}\sum_{k=1}^{|W\cap A_i|}\frac{1}{k}
\]
over the full feasibility family \(F\). Phragmén’s Sequential Rule is generalized through a continuous-load process in which voters earn budget at rate \(1\), candidates have price \(1\), and a candidate is purchased when its supporters collectively accumulate the price. Stable-priceability is generalized through candidate prices \(\pi \in \mathbb{R}_+^C\), unit-budget payments \(p_i:C\to\mathbb{R}_+\), support-only payments, exact coverage of selected items, no profitable deviation to an unselected item, and producer-stability.

The main structural result is exact. PAV satisfies EJR for all elections with matroid constraints, and for any non-matroid \(F\) there exists an election where PAV fails BEJR. Phragmén’s sequential rule satisfies PJR for elections with matroid constraints, and for any non-matroid \(F\) there exists an election where it fails BPJR. For stable-priceability, every stable-priceable outcome satisfies EJR under matroid constraints; if candidate prices are equal, then any stable-priceable outcome satisfies EJR under arbitrary constraints. This identifies matroid feasibility as the boundary at which these strong subset-based proportionality guarantees are available.

The framework recovers familiar settings as special cases. For \(F=\{W\subseteq C:|W|=k\}\), BEJR and EJR reduce to classic multiwinner proportionality. For public decisions with binary issues, \(C\) is partitioned into pairs and exactly one option per issue must be chosen. Diversity-constrained elections with disjoint attribute groups and per-group quotas form a matroid. Ranking and judgment-aggregation style constraints are generally non-matroid, and the necessity results explain why adapted PAV and Phragmén can fail there. The broader significance is that proportional fairness is no longer tied to a fixed-seat committee model; it becomes a property of arbitrary feasible subset selection.

## 3. Swarm-robotic SubCDM

In swarm robotics, SubCDM denotes a decentralized framework in which only a subset of robots performs the decision-making task. The studied problem is best-of-2 decision-making: determining the dominant environmental feature, black versus white tiles. The motivation is explicitly resource-oriented: fewer robots need to sense, move, and communicate; idle robots can be reallocated to other tasks; and performance stagnation or degradation in very large swarms can be avoided [2508.00467].

The evaluated system uses \(100\) foot-bot robots in an \(8 \times 8\) m arena with randomly distributed \(20 \times 20\) cm black and white tiles. Communication is local range-and-bearing with \(d_{\mathrm{comm}}=1\) m, updates are asynchronous at \(10\) ticks/s, and task difficulty is varied by black-tile proportion from \(34\%\) to \(48\%\), corresponding to black:white ratios \(0.52\)–\(0.92\).

SubCDM operates in three phases: subset construction, collective decision-making using DMVD, and subset evaluation or adjustment. Role tenure is stabilized by sampling \(\tau_{\mathrm{role}} \sim \mathrm{Exp}(\text{mean }20\text{ s})\). Two local-information subset-construction strategies are studied. In the leader-based strategy, robots maintain shortest hop count to a leader via
\[
h_i \leftarrow \min(h_i,h_j+1),
\]
and the decision-making subset is
\[
S_{DM}=\{r_i \mid h_i \le s\}.
\]
In the distributed strategy, each robot maintains a local subset parameter \(s_i\) and joins with probability
\[
p_i=0.1 \cdot s_i,\qquad r_i \in S_{DM} \iff \delta_i \le p_i,\ \delta_i \sim U(0,1).
\]
Idle robots can relay up to three messages to maintain connectivity among randomly selected decision-makers.

Subset size is adaptive. In the leader-based variant, the leader waits until at least \(n_{op}=10\) opinions are collected; if the majority ratio exceeds \(r_{op}=0.8\) for \(\tau_{op}=10\) s, the decision at the current \(s\) is recorded, and the process continues until \(k=2\) consistent decisions are obtained. If no stable majority is achieved within \(150\) s, \(s\) is increased. In the distributed variant, each robot tracks confidence \(\alpha_i \in [0,1]\), initialized at \(1\), and updates it when hearing a neighboring decision-maker:
\[
\alpha_i(t+1)=
\begin{cases}
\alpha_i(t)+\gamma(1-\alpha_i(t)), & o_i=o_j\\
\alpha_i(t)-\gamma\alpha_i(t), & o_i\neq o_j
\end{cases}
\]
with \(\gamma=0.01\). Every \(0.05\) decrement in \(\alpha_i\) increments \(s_i\), thereby increasing \(p_i\).

The consensus protocol within \(S_{DM}\) is DMVD. Exploration durations satisfy \(t_e \sim \mathrm{Exp}(\sigma)\) with \(\sigma=10\) s; a robot estimates local quality by \(\hat{\rho}=t_o/t_e\), then disseminates for \(t_d \sim \mathrm{Exp}(\hat{\rho}g)\) with \(g=10\) s. Positive feedback arises because longer dissemination is associated with larger local quality estimates.

Simulation results show that both SubCDM variants maintain accuracy comparable to full-swarm DMVD while using fewer robots. The leader-based subset expands outward and plateaus around \(s=9\) in the \(100\)-robot setup. Spatial organization differs sharply: Moran’s Index over \(100\) runs is \(0.875\) for leader-based and \(0.366\) for distributed selection. Convergence time increases with task difficulty for all methods; full-swarm DMVD is faster, whereas SubCDM trades speed for resource efficiency. Under reduced communication and faults, the distributed variant is more resilient, while the leader-based variant is sensitive to hierarchy disruption and leader faults. The paper explicitly does not provide formal proofs or bounds for convergence rates, error probabilities, or connectivity conditions; support is empirical via repeated ARGoS simulations and stability thresholds.

## 4. Partial-information and communication-complexity formulations

A different line of work interprets subset-based collective decision-making as collective computation under partial information. In the macroscope model, the global input is \(x=(x_1,\dots,x_n)\), the parties are \(P=\{1,\dots,k\}\), and each party \(i\) observes a subset \(S_i \subseteq [n]\). The allotment structure is \(A=\{S_1,\dots,S_k\}\), and a macroscope is the pair \((f,A)\), where \(f:X^n\to Y\) is the global function to be computed [1204.3860].

The model distinguishes two meta-information regimes. In the single-blind regime, each party knows the full allotment structure \(A\). In the double-blind regime, each party knows only its own subset indices and values. Communication is one-round simultaneous broadcast on a blackboard, and the cost is total bits transmitted. This isolates how subset overlap and knowledge of overlap affect the complexity of collective decisions.

General bounds are sharp. Every single-blind macroscope on \(n\) bits has a protocol of cost \(n\), and this is optimal. Every \(k\)-player double-blind macroscope on \(n\) bits has a protocol with cost at most \(2nk\). The single-blind upper bound is achieved by responsibility assignment: for each index \(j\), the lowest-index party holding \(j\) broadcasts \(x_j\). The double-blind upper bound requires each party to reveal both its held indices and the values it sees.

For specific tasks, subset structure can lower communication dramatically. For \(D\)-ary constancy detection, if \(G_A\) is the intersection graph over parties and \(r\) is its number of connected components, then
\[
C_{\text{single-blind}}(\mathrm{Constancy}_D,A)\le r\lceil \log D\rceil + k,
\]
optimal up to factor \(2\). In the double-blind regime,
\[
C_{\text{double-blind}}(\mathrm{Constancy}_D,A)\le k\lceil \log(D+1)\rceil.
\]
For Boolean step-function detection,
\[
C_{\text{double-blind}}(\mathrm{BSF},A)\le 2k\lceil \log n\rceil,\qquad
C_{\text{single-blind}}(\mathrm{BSF},A)\le k\lceil \log n\rceil + 2k.
\]
For approximate averaging with error tolerance \(\epsilon\), single-blind knowledge of duplication counts \(N_j\) yields
\[
C_{\text{single-blind}}(\mathrm{Average}_\epsilon,A)\le k\lceil \log(k/\epsilon)\rceil,
\]
while there exist \(2\)-player double-blind instances with
\[
C_{\text{double-blind}}(\mathrm{Average}_\epsilon,A)\ge n\cdot \log(1/(n\epsilon)).
\]

The central lesson is that subset overlap helps only when it is known. Connected overlaps can compress communication to component summaries, and knowledge of duplication factors prevents double counting. Without such meta-information, the same subset structure becomes a source of uncertainty that must be communicated away.

## 5. Stochastic, higher-order, and neighborhood-based dynamics

Several models treat subset-based collective decisions as emergent stochastic or nonlinear dynamics rather than explicit optimization. In an urn-based process for social choice, an urn contains balls labeled by alternatives, a random voter compares two sampled labels, the preferred label replaces the losing one, and with probability \(r\) a random mutation relabels a ball uniformly. The urn state \(X^{(N,r)}(n)\) is a Markov chain on the discrete simplex, and the expected vector field is
\[
f_i^{(r)}(p)=2(1-r)p_i(\tilde M_R p)_i + r(1/d-p_i),
\]
where \(\tilde M_R\) is the skew-symmetric majority-margin matrix. The main theorem states that for sufficiently small \(r>0\) and large enough \(N\), the urn state spends at least a \(1-\tau\) fraction of time within \(\delta\) of some maximal lottery \(p^\ast\), and the probability of being within \(\delta\) increases exponentially in \(n\) [2103.14351].

A different higher-order formulation models opinion dynamics on hypernetworks. Agents have states \(x_i\), pairwise interactions are encoded by \(A^{(2)}\), and \(3\)-way interactions by \(A^{(3)}\). For \(h=2\),
\[
\dot{x}_i = -\delta_i x_i + \pi\sum_j a^{(2)}_{ij}\psi_j(x_j) + \pi\sum_{j,k} a^{(3)}_{ijk}\psi_j(x_j)\psi_k(x_k).
\]
Here \(\pi\) is a social-effort bifurcation parameter. Without higher-order terms, the system undergoes a symmetric pitchfork at
\[
\pi_1 = \frac{1}{\lambda_n(\Delta^{-1}A^{(2)})}.
\]
With \(3\)-way interactions, Lyapunov–Schmidt reduction yields
\[
\dot{y} = (\pi-\pi_1)y + \kappa_1 y^3 + \kappa_2 y^2,\qquad \kappa_1<0,\ \kappa_2>0,
\]
so the pitchfork is unfolded into a saddle-node plus pitchfork, producing a bistable interval \((\pi_1^\ast,\pi_1)\). In that interval, the community may remain in deadlock or converge to a nontrivial decision depending on initial conditions [2509.05182].

A third model studies well-mixed binary-opinion swarms in which an agent updates after consulting a neighborhood subset of odd size \(G\), sampled without replacement. If \(m\) of \(N\) agents hold opinion \(X_1\), then the number \(k\) of \(X_1\) agents in the sampled subset is hypergeometric:
\[
P(k\mid N,m,G)=\frac{\binom{m}{k}\binom{N-m}{G-k}}{\binom{N}{G}}.
\]
With majority/minority rule weights \(w_k \in \{+1,-1\}\) and noise level \(\epsilon\), the macroscopic drift in \(z=2x-1\) is
\[
\dot z = -\epsilon z + \sum_{k=1}^{G-1} w_k P(k\mid N,m,G).
\]
For majority-dominated rule sets, the dynamics exhibit stable fixed points near \(z=\pm 1\) without noise and interior stable points when noise is added; for all-minority rules, \(z=0\) becomes the single stable fixed point [1410.5738].

Taken together, these models show that subset interactions are not merely a sparsified version of pairwise averaging. They can induce maximal-lottery approximation, exact finite-population drift laws, or bifurcation changes that are absent in pairwise-only systems. This suggests that the combinatorics of which subsets interact is itself a control parameter of collective decision formation.

## 6. Preference learning, optimization, and institutional design

SubCDM also appears in methods that learn preferences over subsets, optimize feasible subset outcomes, or deliberately choose a deciding subset. In robust ordinal regression for subset comparisons, a subset \(S \subseteq N\) is represented by its indicator vector and evaluated by an interaction-aware utility
\[
U(S;\theta,v)=\sum_{T\in \theta} I_S(T)\,v_T,
\]
where \(\theta \subseteq 2^N\) is a learned support of interacting coalitions. Preference data induce a polyhedron \(V_\theta^R\), and robust dominance requires unanimity across all simplest supports and all consistent parameter vectors. Degree minimization is polynomial via the kernel
\[
K^{(k)}(X,Y)=\sum_{i=1}^k \binom{|X\cap Y|}{i},
\]
whereas cardinality, weighted-size, and lexicographic support minimization are NP-hard and addressed by mixed-integer programming. On the IMDb data reported in the paper, ORD attains Prediction Rate \(0.60\), Precision \(0.76\), Recall \(0.83\), and \(F1=0.81\), versus \(F1\) values of \(0.69\) for LR, \(0.70\) for SVM, and \(0.59\) for KNN [2308.03376].

A complementary optimization framework represents collective discrete optimisation as judgment aggregation with weighted issues. An agenda \(\Phi=\{\phi_1,\dots,\phi_m\}\), integrity constraints \(\Gamma\), agent ballots \(B_i \in \{0,1\}^m\), and a set scoring function \(s(B_i,X)\) define modular rules
\[
R_{\mathrm{sum}}^s(B,\Gamma)=\arg\max_{X\in \mathrm{Mod}(\Gamma)} \sum_{i\in N} s(B_i,X),
\]
\[
R_{\mathrm{egal}}^s(B,\Gamma)=\arg\max_{X\in \mathrm{Mod}(\Gamma)} \min_{i\in N} s(B_i,X),
\]
and ranked variants. This framework subsumes approval-based participatory budgeting, collective spanning trees, collective scheduling, and multiwinner rules. The paper proves, among other equivalences, that the median rule is equivalent to \(R_{\mathrm{sum}}^{s^{\mathrm{simp}}}\), the weighted median rule is equivalent to \(R_{\mathrm{sum}}^{s^{\mathrm{wt}}}\), and \(R_{\mathrm{rank}}^{s^{\mathrm{simp}}}\) is equivalent to Ranked Agenda. It also gives an ILP implementation, including a single-commodity-flow formulation for spanning trees [2112.00574].

Institutional SubCDM appears explicitly in the problem of appointing a subset of agents to vote on behalf of the whole group under uncertainty. In the Denying Access Problem (DAP), one removes at most \(\kappa\) agents so that the remaining subset’s majority decisions are guaranteed to coincide with the objectively correct majority outcomes of the full group on as many issues as possible. DAP is NP-complete even in the one-dimensional proposal space, yet fixed-parameter tractable in both \(n\) and \(m\); the \(m\)-parameterized result uses ILP over agent types in \(\{+,-,?\}^m\). In a radical one-dimensional domain, the paper gives polynomial-time algorithms for subset appointment, education, and constrained delegation, showing that endogenous guru sets under delegation can function as effective deciding subsets [2207.04983].

Group design for multidimensional decisions adds a further institutional layer. A complex problem is decomposed into \(M\) independent sub-problems, group members have competence matrix \(C=[c_{i,j}] \in [0,1]^{N\times M}\), proposal quality is \(Q_{i,j}=c_{i,j}\), and evaluations satisfy
\[
E_{i,j}^{i'} = Q_{i,j} c_{i',j} + (1-c_{i',j})\mathrm{Rand}.
\]
Fitness is \(F=Q-C\), with competence cost \(\langle c_0 c_{i,j}^e\rangle\). The main qualitative result is that the best performing groups have at least one specialist for each sub-problem, but specialists also need some insight into the other sub-problems. Empirical analysis on ISI Web of Science data reports that, across nine trends, median citations increase with authors’ average interdisciplinarity, measured by the Shannon entropy of subject classes in their references [1612.00071].

These approaches are methodologically diverse, but they share a design principle: subset structure can be learned, optimized, or institutionally imposed rather than passively inherited.

## 7. Limitations and open directions

The main limitations are domain-specific and technically consequential. In proportional social choice, the positive characterizations for adapted PAV and Phragmén stop exactly at matroid feasibility; under non-matroid constraints, BEJR or PJR can fail, and weighted-candidate settings such as participatory budgeting remain problematic for exact EJR/PJR guarantees [2307.06077]. In swarm robotics, convergence and robustness claims are empirical rather than theorem-level, and the leader-based variant is sensitive to reduced communication and leader faults, while the distributed variant depends on relay forwarding with a cap of three messages per relay [2508.00467].

For partial-information macroscopes, the one-round deterministic model leaves open the effect of multiple rounds, graph-constrained communication, randomized protocols, and intermediate meta-information regimes between single-blind and double-blind [1204.3860]. In higher-order dynamics, heterogeneous nonlinearities, antagonistic interactions, noise, asynchronous updates, time-varying hypernetworks, and hyperedges of order greater than three are identified as extensions that may alter stability and bifurcation structure [2509.05182]. In robust ordinal regression, lexicographic sparsity selection is NP-hard and robust unanimity is intentionally conservative, often yielding abstention when the data do not support reliable prediction [2308.03376]. In collective discrete optimisation as judgment aggregation, modular ILP formulations are expressive but can be outperformed by specialized algorithms, and strategyproofness and general proportionality notions are not resolved [2112.00574].

A final conceptual point follows from the literature as a whole. SubCDM is best understood not as a single algorithmic family but as a structural idea: collective decisions can be mediated by subsets, and the mathematical consequences depend on what those subsets represent—feasible outcomes, cohesive coalitions, active robots, sampled alternatives, observed indices, or specialized task domains. The technical results surveyed here show that subset structure can yield stronger fairness axioms, lower communication, lower resource use, richer nonlinear dynamics, or more reliable preference inference, but only under correspondingly specific assumptions about feasibility, information, or interaction topology.

Source: https://www.emergentmind.com/topics/subset-based-collective-decision-making-subcdm