Collective Recourse
- Collective recourse is an extension of individual recourse that models population-level impacts, fairness constraints, and multi-agent interactions.
- It integrates external costs, performative dynamics, and strategic behaviors to address systemic effects beyond minimal counterfactual changes.
- Methodologies like capacitated matching, population summaries (AReS), and online collective action illustrate trade-offs between individual benefits and overall societal welfare.
Searching arXiv for the cited papers and closely related work on collective or multi-agent algorithmic recourse. arxiv_search(query="collective recourse algorithmic recourse multi-agent performative validity endogenous macrodynamics AReS RAGUEL", max_results=10) arxiv_search(query="collective recourse algorithmic recourse", max_results=10) search_arxiv("collective recourse algorithmic recourse") Collective recourse denotes a family of extensions of algorithmic recourse in which the unit of analysis is no longer an isolated individual interacting with a fixed model. In this broader literature, the term covers at least four related settings: the joint effects that arise when many individuals implement recourse and thereby shift the data distribution or the model itself; multi-agent environments in which recourse recommendations affect third parties, provider capacities, or strategic interaction; global or group-level summaries of recourse across a population; and socio-technical pipelines for correcting recurring group harms in deployed generative systems. Across these settings, the common move is from minimal-cost counterfactual change for one instance to system-level design, welfare, validity, governance, and fairness under collective action (Altmeyer et al., 2023).
1. Conceptual scope
The single-individual recourse problem is typically formulated as follows: for a pre-trained model and a negative-outcome instance , find the minimal-cost perturbation so that , or, equivalently, solve
This formulation appears explicitly in work on endogenous macrodynamics, where it serves as the baseline from which collective recourse departs (Altmeyer et al., 2023).
In the literature, collective recourse is not a single formalism. In “Beyond Individualized Recourse: Interpretable and Interactive Summaries of Actionable Recourses” (Rawal et al., 2020), it means a small, human-readable summary that prescribes recourses for most affected individuals by means of recourse rules over subpopulations. In “Multi-Agent Algorithmic Recourse” (O'Brien et al., 2021), it refers to a multi-agent environment in which current approaches to algorithmic recourse fail to guarantee certain ethically desirable properties once the assumption of a single agent environment is relaxed. In “Endogenous Macrodynamics in Algorithmic Recourse” (Altmeyer et al., 2023), it denotes the inclusion of an external cost of recourse, reflecting negative spill-over on the remainder of the population. In “Online Algorithmic Recourse by Collective Action” (Creager et al., 2023), it denotes the fact that a crowd of users can collectively shape future predictions for a target individual by leveraging the parameter update rule. In “Collective Recourse for Generative Urban Visualizations” (Mushkani, 15 Sep 2025), it is formalized as structured community “visual bug reports” that trigger fixes to models and planning workflows.
These formulations share a negative result about individualized analysis. Past work has largely focused on the effect algorithmic recourse has on a single agent; many methodologies assume a static environment; and several papers argue that these assumptions suppress hidden externalities, feedback loops, or group-level harms. This suggests that “collective recourse” is best understood as an umbrella term for recourse methods that explicitly model population-level consequences, multi-stakeholder constraints, or recurring harms rather than only individual actionability.
2. Formalizations beyond the individual instance
One formalization augments the standard recourse objective with an externality term. In “Endogenous Macrodynamics in Algorithmic Recourse” (Altmeyer et al., 2023), the unified objective is
Here is a possibly latent search state, maps back to the feature space, and trade-off private versus external costs. Wachter’s method sets 0, 1, 2; latent-space methods such as REVISE and CLUE learn 3 via a VAE decoder; and DiCE adds a Determinantal Point Process term in 4 to enforce diversity (Altmeyer et al., 2023).
A second formalization treats collective recourse as population-level summarization. AReS represents collective recourse as a two-level recourse set
5
where 6 is a Boolean conjunction of predicates defining a subpopulation, and 7 is an unordered set of inner recourse rules. The framework quantifies recourse correctness, coverage, cost through 8 and 9, and interpretability through 0, 1, and 2, then solves
3
Theorem 2 states that if 4, 5, 6, and 7, then AReS reduces exactly to individual-recouse objectives used by methods such as Wachter et al. (2018), Ustun et al. (2019), and MACE (Rawal et al., 2020).
A third formalization is explicitly many-to-many. In “From Individual to Multi-Agent Algorithmic Recourse: Minimizing the Welfare Gap via Capacitated Bipartite Matching” (Khotanlou et al., 14 Aug 2025), there is a set of seekers 8 and a set of providers 9, each provider 0 has finite capacity 1, each seeker 2 has precomputed pairwise costs 3, and the central planner coordinates matches to maximize total benefit subject to capacities. Costs are transformed to weights, for example
4
and the basic capacitated matching is
5
subject to each seeker matched at most once and each provider respecting capacity. The welfare gap is defined as
6
measuring loss due to competition for limited provider capacity (Khotanlou et al., 14 Aug 2025).
A fourth formalization is socio-technical rather than optimization-centric. In generative urban visualization, a visual bug report is
7
with 8 the affected group, 9 the urban context, 0 severity, 1 the number of unique reporters, 2 representativeness, and 3 evidence quality. Its mandate score is
4
and a report mandates a fix when 5 (Mushkani, 15 Sep 2025).
3. Endogeneity, performativity, and hidden external cost
A central claim in the collective-recouse literature is that recourse is endogenous. In the simulation framework of “Endogenous Macrodynamics in Algorithmic Recourse” (Altmeyer et al., 2023), at each round 6 there is a population distribution 7 over 8; a batch 9 is sampled from the negative-class portion 0; factuals 1 are replaced with their counterfactuals 2 via generator 3; and the model is retrained or fine-tuned, 4, on the updated data, thereby inducing 5. The paper reports domain shifts as counterfactuals accumulate, model shifts as retraining changes parameters 6 and the decision boundary, and macrodynamics in which repeated recourse leads to cumulative 7-shift, cumulative model-drift, and possible erosion of accuracy for future applicants. All generators induce statistically significant domain and model shifts, and model accuracy 8 typically drops 5–15 ppt after 9 rounds with 5% of negatives per round (Altmeyer et al., 2023).
The paper operationalizes external cost as the negative effect on the remaining population, for example lower accuracy for future applicants or bank losses when admitting less-creditworthy borrowers. Quantification is proposed in three forms: the aggregate increase in MMD over time, the cumulative drop in 0-score of 1 over 2, and the sum of pairwise model Disagreement over all retraining steps. This makes precise the claim that any choice of recourse that minimizes only 3 ignores 4 and can lead to socially inefficient outcomes (Altmeyer et al., 2023).
A related but distinct line of work characterizes recourse as performative. In “Performative Validity of Recourse Explanations” (König et al., 18 Jun 2025), recourse explanations are performative because, when many applicants act according to their recommendations, their collective behavior may change statistical regularities in the data and, once the model is refitted, also the decision boundary. Performative validity is defined on a region 5 by requiring that every 6 that was reversed by the original model is also accepted by the retrained model. Proposition 4.1 states that the original predictor 7 and retrained predictor 8 agree at 9 iff under the post-recourse distribution the recourse action 0 is uninformative about the post-recourse label: 1 Theorem 4.2 identifies two sources of invalidity: recourse is influenced by effect variables 2, or recourse intervenes on effect variables 3. Corollary 4.4 states that, under either Assumption 3.1 or Assumption 3.2, any recourse method that abstains from intervening on direct effects is performatively valid; in particular, Improvement-focused Causal Recourse is the only standard method with this guarantee, while CE and CR may fail (König et al., 18 Jun 2025).
The empirical evidence in that paper is sharper than a generic warning. After simulating recourse by sampling from 4 and retraining on a 2:1 mixture of pre- and post-recourse data 5, the authors measure the shift in 6 and the drop in acceptance rate among recourse-implementers. CE and CR induce large negative shifts, up to 7 in conditional probability, and acceptance-rate drops up to 8, while ICR remains stable with zero shift and zero drop across all SCMs (König et al., 18 Jun 2025).
These results make a common misconception untenable: that a recourse explanation remains valid simply because it flips the current predictor for one instance. The collective and performative analyses show that validity can fail after deployment-scale uptake, even when the original recommendation was individually valid.
4. Multi-agent interaction, welfare, and collective action
In explicit strategic settings, collective recourse is tied to third-party effects and welfare constraints. “Multi-Agent Algorithmic Recourse” (O'Brien et al., 2021) studies environments in which recommendations to one player may harm others or reduce total group payoff. In the follow-on study summarized in the details, players were randomly paired in an iterated Prisoner’s Dilemma with an indefinite horizon, with probability 9 the next round occurs and otherwise the game ends; 0. A control group with identical expected length used fixed-length games of exactly 1, 2, or 4 rounds respectively. In total 1 games satisfied 2 or 3 with matching fixed controls. For each single-round subgame, the system pretended to advise one principle player using three optimization variants: single-agent optimal recourse, Pareto-efficient recourse, and social-welfare-efficient recourse (O'Brien et al., 2021).
The key findings are directly relevant to collective recourse. Under single-agent optimal recourse, 434 recommendations would have improved the principle player’s expected payoff but simultaneously reduced total social welfare and harmed the other player, violating Pareto efficiency. Zero recommendations ever improved both the focal player and the other player or increased social welfare. Under Pareto-efficient or joint constraints, no recommendations were possible that both helped the principle player and satisfied these group-level constraints; in effect, the framework would have abstained from advising any strategy change. Under social-welfare-only constraints, 2 860 recommendations would have increased the sum of both players’ payoffs, but all of these would have decreased the focal player’s payoff, and none would have helped the principle player (O'Brien et al., 2021).
This establishes a recurring ethical dilemma in collective recourse: should an algorithm advise an individual to do something that makes them worse off in order to improve overall social welfare? The paper’s conclusion is that single-agent recourse can repeatedly suggest “defect” even though this systematically erodes mutual payoff, whereas imposing Pareto or social-welfare constraints entirely eliminates harmful recommendations at the cost of sometimes giving no advice (O'Brien et al., 2021).
A different multi-agent formulation replaces strategic interaction with capacity-constrained providers. In the capacitated bipartite matching framework, three layers are defined: basic capacitated matching with fixed capacities, optimal capacity redistribution with fixed total capacity 4, and cost-aware optimization with penalties 5 for capacity changes. The heuristic for capacity redistribution sorts the top-6 individual matches 7 and sets 8 equal to the number of these matches pointing to 9; it runs in 0 and is provably optimal for Problem (2) under the top-1 criterion (Khotanlou et al., 14 Aug 2025).
The experimental summary reports that, on Two-Moon, basic matching with capacity 2 achieves social welfare 5.59, or 93.1% of the individual-welfare upper bound; optimal redistribution with capacity 3 achieves 6.01, or 100.0%; and cost-aware optimization with capacity 4 achieves 5.97, or 99.4%. Across Credit and COMPAS, basic matching attains approximately 94–96% of IW, redistribution yields 100%, and cost-aware retains approximately 99% while moderating capacity changes (Khotanlou et al., 14 Aug 2025). A plausible implication is that “collective recourse” in this line of work is not merely about explaining rejected decisions but about redesigning allocation mechanisms so that individual actionability remains compatible with collectively feasible outcomes.
A third variant studies collective action in online learning. In “Online Algorithmic Recourse by Collective Action” (Creager et al., 2023), online algorithmic recourse is formalized as a bi-level problem in which all data subjects may perturb their own inputs 5 so as to shape the updated 6 and obtain a goal class for a query subject: 7 subject to
8
On UCI-Iris and MNIST, collective recourse is more effective than individual recourse at a given 9 budget; the reported qualitative result is that collective always achieves lower loss for any 00 in Figure 1 (Creager et al., 2023). The paper does not provide formal theorems, but its main theoretical observation is that by violating the i.i.d. assumption, coordinating perturbations allows data subjects to gain leverage over model behavior that individual recourse alone cannot.
5. Group-level summaries, ranking fairness, and operational mechanisms
Not all collective-recouse frameworks are about feedback loops or games. Some operate at the level of explanation, audit, and ranking. AReS is explicitly designed to construct global counterfactual explanations that provide an interpretable and accurate summary of recourses for the entire population. Its objective simultaneously optimizes for correctness of recourses and interpretability of explanations while minimizing overall recourse costs across the entire population. The framework yields a small number of compact rule sets capturing recourses for well defined subpopulations, and Theorem 3 states that if the classifier admits recourse for every affected individual, then
01
where 02 is the approximation ratio of the submodular-maximization algorithm (Rawal et al., 2020).
The empirical summary for AReS emphasizes interpretability and auditability. On credit scoring, bail decisions, and COMPAS recidivism, AReS matches or exceeds state-of-the-art individualized methods in recourse accuracy, often 03, while reducing mean cost by 20–50%. On a 3-layer DNN for COMPAS, AReS achieves 99.4% accuracy at mean cost 2.8, versus 73.7% and cost 4.5 for FACE. In a user study with a hand-crafted biased bail black box, users shown AReS summaries immediately reported the asymmetric treatment, with more than 90% detecting bias, whereas aggregated individual recourses confused 80% of users (Rawal et al., 2020). These results support a distinct interpretation of collective recourse: a global explanatory summary that reveals subgroup disparities invisible when inspecting individuals one at a time.
RAGUEL introduces another collective notion: ranked group-level recourse fairness. Let 04 be partitioned into disjoint subgroups 05, let 06 be the global proportion of the protected subgroup, and define subgroup mean recourse cost
07
Then the subgroup-cost ratio
08
measures cost balance, with 09 perfect balance. A set satisfies 10-group-recourse fairness if 11, and an ordered prefix is ranked-recourse-fair if every prefix satisfies both 12-fair representation and 13-group-recourse fairness (Haldar et al., 2022).
The recourse cost model uses weighted Euclidean distance to the nearest counterfactual: 14 RAGUEL then re-ranks while minimizing total perturbation, and, in the block-based version, enforces constraints within each block. The empirical results on DC3 report an initial 15; post-process classifiers may worsen 16 to 0.472 or 0.422; FA*IR at 400 items achieves 17; FoEiR at 50 items gets approximately 0.793; and RAGUEL raises 18 to 0.876, with the block version giving 19 with approximately 130% lower total cost of modifications. On counterfactual generation, RAGUEL-CF attains closeness approximately 1.4, sparsity 1.36 dimensions, and time 0.003 s, compared with MACE at closeness approximately 3 095, sparsity 2.23, time 1.7 s, and AR at closeness approximately 9.5, sparsity 1.6, time 0.003 s (Haldar et al., 2022).
A final operational mechanism appears in generative urban visualization. There, collective recourse is a five-stage pipeline: Report, Triage, Fix, Verify, Closure. Triage clusters reports by harm type and context and computes 20; only reports with 21 enter the fix stage. Fix selection uses a quick steerability test on the subgroup visual eval set (SVES): if targeted counter- or negative prompts reduce recurrence on SVES, choose them as first response; else if the root cause is a data coverage gap, choose dataset edit; else choose reward-model tweak. Verification re-generates the SVES pre- and post-fix, checks that recurrence falls below a harm-specific tolerance, and collects subgroup satisfaction; closure requires rotating citizen jury approval and updates public dashboards, model cards, datasheets, and planning artifacts (Mushkani, 15 Sep 2025).
6. Trade-offs, safeguards, and unresolved tensions
Across the literature, collective recourse introduces explicit trade-offs rather than eliminating them. In the multi-agent Prisoner’s Dilemma setting, single-agent optimal recourse helps the principle player in 434 cases while harming the other player and reducing total welfare; Pareto or social-welfare constraints prevent such harm but may force abstention; and social-welfare-only constraints generate 2 860 recommendations that improve the sum of payoffs while making the focal player worse off (O'Brien et al., 2021). In the capacitated matching setting, redistribution can close the welfare gap, but cost-aware optimization is introduced precisely because changing capacities may itself be costly (Khotanlou et al., 14 Aug 2025).
In endogenous and performative settings, the main tension is between private ease of recourse and societal stability. “Endogenous Macrodynamics in Algorithmic Recourse” (Altmeyer et al., 2023) proposes several mitigation strategies that are special cases of the collective objective: a more conservative decision threshold 22, Classifier-Preserving ROAR with
23
and Gravitational Counterfactuals with
24
where 25 is the mean of positive-class points. Combinations such as ClaPROAR+LS search and higher 26 plus gravitational counterfactuals yield near-zero domain drift and much reduced model-drift in the reported figures (Altmeyer et al., 2023). “Performative Validity of Recourse Explanations” (König et al., 18 Jun 2025) gives a sharper design rule: constrain recourse suggestions to causal features of the target variable and abstain from intervening on direct effects.
In governance-oriented settings, the trade-offs are procedural. The urban-visualization framework defines mandate thresholds by target precision–recall trade-offs: at 27, flagged share is 73.8%, precision 83.6%, recall 94.9%; at 28, flagged share is 52.5%, precision 92.9%, recall 75.0%; at 29, precision rises to 96.4% while recall falls to 51.9%; and at 30, precision is 100% with recall 31.4%. Increasing representativeness by 31 raises recall at 32 from 75.0% to 81.4% while precision dips from 92.9% to 91.4% (Mushkani, 15 Sep 2025). The paper explicitly identifies risks: overfitting to vocal minorities, astroturfing or gaming the system, and a transparency trap in which dashboards exist without action. Its safeguards are rate-limiting, deduplication, evidence-quality checks, safe-harbor protections, public dashboards with stratified service-level agreements, rotating mini-publics, and separation of ideation renderings from evidentiary bug reports (Mushkani, 15 Sep 2025).
A recurrent misconception is that collective recourse is simply “individual recourse applied many times.” The surveyed work contradicts that view. Once recommendations interact through retraining, ranking, capacities, strategic behavior, or community reporting, new objects appear: external cost, welfare gap, representativeness, mandate thresholds, ranked group-level recourse fairness, and performative validity. A plausible implication is that collective recourse is less a single algorithm than a design stance: recourse must be evaluated with respect to the population, the update rule, the allocation mechanism, and the institutional process within which recommendations are acted upon.