Papers
Topics
Authors
Recent
Search
2000 character limit reached

Add-Remove Heterogeneous Differential Privacy (AHDP)

Updated 10 July 2026
  • AHDP is a set of differential privacy formulations that replace a global budget with structured, context-aware guarantees via add/remove neighboring datasets.
  • It encompasses diverse models including graph-based, coordinate-wise, and correlation-aware settings, each employing unique optimization techniques and trade-offs.
  • AHDP frameworks enable practical implementations with efficient algorithms, offering improved utility and tailored privacy protection under varying demands.

Add-remove Heterogeneous Differential Privacy (AHDP) denotes a family of differential privacy formulations that retain add/remove-style neighboring datasets while allowing privacy protection to vary across users, dataset pairs, or typed records. In the cited literature, the term appears in closely related but not identical formal settings: graph-based heterogeneous DP with edge-dependent privacy parameters and utility-optimal extension from a boundary set (Torkamani et al., 2022); coordinate-wise heterogeneous privacy demands for mean estimation under add/remove adjacency (Chaudhuri et al., 2023); and a correlation-aware framework in which each record is a pair (x,ϵ)(x,\epsilon) and privacy is measured by a weighted add–remove distance that jointly accounts for data and privacy demand (Chaudhuri et al., 2 Sep 2025). Across these formulations, the common objective is to replace a single global privacy budget with structured heterogeneity while preserving operational DP guarantees.

1. Formal scope and principal formulations

The literature uses AHDP to describe heterogeneous privacy under add/remove semantics, but the primitive object that carries heterogeneity differs by paper. In one line of work, datasets are vertices of a graph and heterogeneity is attached to edges through a symmetric function ε:E[0,)\varepsilon:E\to[0,\infty) (Torkamani et al., 2022). In another, datasets are vectors XXn\mathbf X\in\mathcal X^n and user ii carries an individual budget ϵi\epsilon_i that governs all neighboring dataset pairs differing only in coordinate ii (Chaudhuri et al., 2023). In the correlation-aware formulation, each input is a pair (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}, datasets are multisets of such pairs, and privacy is defined through a weighted add–remove metric dαd_\alpha (Chaudhuri et al., 2 Sep 2025).

Formulation Dataset model Heterogeneity object
Graph-based heterogeneous DP Graph G=(V,E)\mathcal G=(V,E) with datasets as vertices Edge budget ε(u,v)\varepsilon(u,v)
Coordinate-wise heterogeneous DP ε:E[0,)\varepsilon:E\to[0,\infty)0 User budget ε:E[0,)\varepsilon:E\to[0,\infty)1
Correlation-aware AHDP Multiset of ε:E[0,)\varepsilon:E\to[0,\infty)2 pairs Weight ε:E[0,)\varepsilon:E\to[0,\infty)3

For the graph-based definition, a randomized mechanism ε:E[0,)\varepsilon:E\to[0,\infty)4 is ε:E[0,)\varepsilon:E\to[0,\infty)5-DP if for every edge ε:E[0,)\varepsilon:E\to[0,\infty)6 and every event ε:E[0,)\varepsilon:E\to[0,\infty)7,

ε:E[0,)\varepsilon:E\to[0,\infty)8

For the coordinate-wise definition, a mechanism ε:E[0,)\varepsilon:E\to[0,\infty)9 satisfies XXn\mathbf X\in\mathcal X^n0-Heterogeneous DP if, for every XXn\mathbf X\in\mathcal X^n1, every pair of datasets differing only in coordinate XXn\mathbf X\in\mathcal X^n2, and every event XXn\mathbf X\in\mathcal X^n3,

XXn\mathbf X\in\mathcal X^n4

For the correlation-aware definition, letting XXn\mathbf X\in\mathcal X^n5 denote the count of type XXn\mathbf X\in\mathcal X^n6 in dataset XXn\mathbf X\in\mathcal X^n7 and fixing XXn\mathbf X\in\mathcal X^n8, the add–remove distance is

XXn\mathbf X\in\mathcal X^n9

and ii0 is ii1-AHDP if for every ii2 and measurable ii3,

ii4

These definitions share add/remove structure but not a single canonical state space. This suggests that AHDP is best understood as a heterogeneous-privacy design pattern rather than a unique formalism.

2. Graph-based AHDP and utility-optimal boundary extension

In "Heterogeneous Differential Privacy via Graphs" (Torkamani et al., 2022), the starting point is a simple undirected graph ii5 in which each vertex is a database and ii6 means that ii7 can be obtained from ii8 by adding or removing exactly one record. For a binary true query ii9, the boundary set ϵi\epsilon_i0 consists of those ϵi\epsilon_i1 for which there exists a neighbor ϵi\epsilon_i2 with ϵi\epsilon_i3. The mechanism is first specified on this boundary, with Bernoulli output probabilities

ϵi\epsilon_i4

and the central problem is to extend ϵi\epsilon_i5 from ϵi\epsilon_i6 to all of ϵi\epsilon_i7 while preserving heterogeneous DP and maximizing pointwise utility.

The key device is the strongest induced DP condition. Although DP is imposed only on edges, path chaining yields induced upper bounds on ϵi\epsilon_i8 from a starting value ϵi\epsilon_i9. Along an edge ii0, the paper records the two inequalities

ii1

and

ii2

Composing these linear maps along a path ii3 gives an induced upper bound ii4, and the strongest induced bound is

ii5

Compatibility of boundary probabilities is then checked by requiring

ii6

for all ii7.

If the boundary assignment is compatible, the optimal extension to any ii8 is

ii9

and

(x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}0

The paper states that this extension is the unique maximizer of (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}1 at each (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}2 subject to all heterogeneous DP constraints. It also gives a path-graph closed form: on a simple path (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}3 with budgets (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}4 and (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}5, the unique optimal extension has a two-regime formula determined by a threshold index (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}6, and the resulting pathwise bias obeys

(x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}7

Computationally, each (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}8 can be computed in polynomial time by a Dijkstra-style shortest-path routine in (x,ϵ)WX×R0(x,\epsilon)\in\mathcal W\subseteq \mathcal X\times \mathbb R_{\ge 0}9. Testing compatibility on the boundary requires dαd_\alpha0, extending to all vertices requires dαd_\alpha1, and the worst-case complexity is dαd_\alpha2. The authors note that the algorithm relies only on shortest-path subroutines and is directly implementable in any graph library; when dαd_\alpha3 is much smaller than dαd_\alpha4, the effective runtime is correspondingly lower (Torkamani et al., 2022).

3. Coordinate-wise heterogeneous privacy demands and AHDP mean estimation

"Mean Estimation Under Heterogeneous Privacy Demands" studies bounded univariate mean estimation under user-specific privacy levels (Chaudhuri et al., 2023). The data domain is dαd_\alpha5, datasets are vectors dαd_\alpha6, and heterogeneous privacy is indexed by dαd_\alpha7, typically ordered so that dαd_\alpha8. The privacy-loss formulation is coordinate specific: dαd_\alpha9

The paper’s concrete mechanism is the Affine Differentially Private Mean (ADPM). After sorting the budgets, it constructs a sequence G=(V,E)\mathcal G=(V,E)0 recursively: G=(V,E)\mathcal G=(V,E)1 where G=(V,E)\mathcal G=(V,E)2 and G=(V,E)\mathcal G=(V,E)3 are updated cumulatively as G=(V,E)\mathcal G=(V,E)4 and G=(V,E)\mathcal G=(V,E)5. If

G=(V,E)\mathcal G=(V,E)6

the mechanism returns the constant estimator G=(V,E)\mathcal G=(V,E)7. Otherwise it sets

G=(V,E)\mathcal G=(V,E)8

draws G=(V,E)\mathcal G=(V,E)9, and outputs

ε(u,v)\varepsilon(u,v)0

The complexity is ε(u,v)\varepsilon(u,v)1 time and ε(u,v)\varepsilon(u,v)2 memory, with near-linear time possible when the ε(u,v)\varepsilon(u,v)3 come from a small discrete set.

The paper ties privacy calibration directly to coordinate sensitivity. For the affine estimator ε(u,v)\varepsilon(u,v)4, the change induced by varying ε(u,v)\varepsilon(u,v)5 over ε(u,v)\varepsilon(u,v)6 is ε(u,v)\varepsilon(u,v)7. A Laplace mechanism with scale ε(u,v)\varepsilon(u,v)8 is ε(u,v)\varepsilon(u,v)9-DP at coordinate ε:E[0,)\varepsilon:E\to[0,\infty)00 if ε:E[0,)\varepsilon:E\to[0,\infty)01, so choosing

ε:E[0,)\varepsilon:E\to[0,\infty)02

suffices. In ADPM, the construction ensures ε:E[0,)\varepsilon:E\to[0,\infty)03, hence ε:E[0,)\varepsilon:E\to[0,\infty)04 enforces all coordinate-wise privacy constraints simultaneously. The summary emphasizes that no “standard” composition theorem is needed because a single global noise scale is calibrated against all user budgets.

The utility analysis is minimax. Defining

ε:E[0,)\varepsilon:E\to[0,\infty)05

the minimax mean-squared error satisfies

ε:E[0,)\varepsilon:E\to[0,\infty)06

for universal constants ε:E[0,)\varepsilon:E\to[0,\infty)07, and ADPM achieves the upper bound with ε:E[0,)\varepsilon:E\to[0,\infty)08 up to small constant factors. The optimization underlying the upper bound is

ε:E[0,)\varepsilon:E\to[0,\infty)09

with a water-filling argument, quasi-convexity, and KKT conditions yielding the recursion implemented by ADPM.

A central structural phenomenon is saturation. If ε:E[0,)\varepsilon:E\to[0,\infty)10 is the smallest index such that

ε:E[0,)\varepsilon:E\to[0,\infty)11

then for ε:E[0,)\varepsilon:E\to[0,\infty)12 one has ε:E[0,)\varepsilon:E\to[0,\infty)13, while for ε:E[0,)\varepsilon:E\to[0,\infty)14 the recursion saturates and all ε:E[0,)\varepsilon:E\to[0,\infty)15 become equal to the same common value independent of ε:E[0,)\varepsilon:E\to[0,\infty)16. The paper interprets this as a “privacy-for-free” effect: once the most stringent users determine the error rate, increasing the budgets of less stringent users does not further improve the minimax rate (Chaudhuri et al., 2023).

4. Correlation-aware AHDP and its operational meaning

"Managing Correlations in Data and Privacy Demand" argues that conventional HDP can fail when user data and privacy demand are correlated (Chaudhuri et al., 2 Sep 2025). In that setting, holding ε:E[0,)\varepsilon:E\to[0,\infty)17 fixed while varying only ε:E[0,)\varepsilon:E\to[0,\infty)18 is insufficient because the pair ε:E[0,)\varepsilon:E\to[0,\infty)19 may itself be private and correlated. The proposed remedy is a correlation-aware AHDP framework in which each tuple ε:E[0,)\varepsilon:E\to[0,\infty)20 is treated as an atomic type and neighboring changes are measured on multisets of such pairs.

Formally, with dataset ε:E[0,)\varepsilon:E\to[0,\infty)21 and count function ε:E[0,)\varepsilon:E\to[0,\infty)22, the weight function ε:E[0,)\varepsilon:E\to[0,\infty)23 is required later to satisfy ε:E[0,)\varepsilon:E\to[0,\infty)24 for all ε:E[0,)\varepsilon:E\to[0,\infty)25. The add–remove distance is

ε:E[0,)\varepsilon:E\to[0,\infty)26

A randomized mechanism is ε:E[0,)\varepsilon:E\to[0,\infty)27-AHDP if

ε:E[0,)\varepsilon:E\to[0,\infty)28

for every ε:E[0,)\varepsilon:E\to[0,\infty)29 and measurable ε:E[0,)\varepsilon:E\to[0,\infty)30. If ε:E[0,)\varepsilon:E\to[0,\infty)31 is homogeneous and ε:E[0,)\varepsilon:E\to[0,\infty)32, this recovers add–remove ε:E[0,)\varepsilon:E\to[0,\infty)33-DP. The summary also states that post-processing and composition hold exactly as in ordinary DP, with ε:E[0,)\varepsilon:E\to[0,\infty)34.

The paper supplies an operational interpretation through binary hypothesis testing. Given observations ε:E[0,)\varepsilon:E\to[0,\infty)35 and hypotheses

ε:E[0,)\varepsilon:E\to[0,\infty)36

for a rejection region ε:E[0,)\varepsilon:E\to[0,\infty)37 define

ε:E[0,)\varepsilon:E\to[0,\infty)38

Then ε:E[0,)\varepsilon:E\to[0,\infty)39 is ε:E[0,)\varepsilon:E\to[0,\infty)40-AHDP if and only if

ε:E[0,)\varepsilon:E\to[0,\infty)41

The summary further states that, in AHDP, if an adversary might be distinguishing ε:E[0,)\varepsilon:E\to[0,\infty)42 from ε:E[0,)\varepsilon:E\to[0,\infty)43, her power is at most ε:E[0,)\varepsilon:E\to[0,\infty)44, and that for a sequence of threat models with ε:E[0,)\varepsilon:E\to[0,\infty)45 unseen points one obtains

ε:E[0,)\varepsilon:E\to[0,\infty)46

where ε:E[0,)\varepsilon:E\to[0,\infty)47.

A major contrast with standard HDP is explicit: the summary states that standard HDP can completely fail under correlations, and that one can leak the entire dataset in some contrived but valid HDP mechanism. The AHDP reformulation is presented as robust to such correlations because the privacy guarantee is indexed by the combined data–privacy type rather than data alone (Chaudhuri et al., 2 Sep 2025).

5. Universal mechanisms under correlation-aware AHDP

The correlation-aware paper develops universal AHDP mechanisms, meaning mechanisms that are ε:E[0,)\varepsilon:E\to[0,\infty)48-AHDP for every finite ε:E[0,)\varepsilon:E\to[0,\infty)49 as soon as the chosen weights satisfy ε:E[0,)\varepsilon:E\to[0,\infty)50 (Chaudhuri et al., 2 Sep 2025). The summary gives four concrete families.

The first is a Sampling Mechanism, attributed to Jørgensen et al. ’15. Given dataset ε:E[0,)\varepsilon:E\to[0,\infty)51, thresholds ε:E[0,)\varepsilon:E\to[0,\infty)52, and ε:E[0,)\varepsilon:E\to[0,\infty)53, each ε:E[0,)\varepsilon:E\to[0,\infty)54 is included in a subsample ε:E[0,)\varepsilon:E\to[0,\infty)55 independently with probability ε:E[0,)\varepsilon:E\to[0,\infty)56, after which any homogeneous ε:E[0,)\varepsilon:E\to[0,\infty)57-DP algorithm is run on ε:E[0,)\varepsilon:E\to[0,\infty)58. The resulting mechanism is stated to be ε:E[0,)\varepsilon:E\to[0,\infty)59-AHDP.

The second family covers linear-counting queries such as sums, counts, and histograms. For ε:E[0,)\varepsilon:E\to[0,\infty)60 with ε:E[0,)\varepsilon:E\to[0,\infty)61 and

ε:E[0,)\varepsilon:E\to[0,\infty)62

the mechanism is

ε:E[0,)\varepsilon:E\to[0,\infty)63

Its add–remove sensitivity in the weighted metric is exactly ε:E[0,)\varepsilon:E\to[0,\infty)64, which justifies the Laplace scale, and any choice ε:E[0,)\varepsilon:E\to[0,\infty)65 yields universal AHDP.

The third family is mean estimation via two queries for ε:E[0,)\varepsilon:E\to[0,\infty)66. With two weight functions ε:E[0,)\varepsilon:E\to[0,\infty)67, the numerator and denominator are

ε:E[0,)\varepsilon:E\to[0,\infty)68

and

ε:E[0,)\varepsilon:E\to[0,\infty)69

and the output is

ε:E[0,)\varepsilon:E\to[0,\infty)70

By composition, ε:E[0,)\varepsilon:E\to[0,\infty)71 is ε:E[0,)\varepsilon:E\to[0,\infty)72-AHDP and ε:E[0,)\varepsilon:E\to[0,\infty)73 is ε:E[0,)\varepsilon:E\to[0,\infty)74-AHDP, so ε:E[0,)\varepsilon:E\to[0,\infty)75 is ε:E[0,)\varepsilon:E\to[0,\infty)76-AHDP.

The fourth family is linear regression. For data

ε:E[0,)\varepsilon:E\to[0,\infty)77

choose weights ε:E[0,)\varepsilon:E\to[0,\infty)78, form ε:E[0,)\varepsilon:E\to[0,\infty)79, ε:E[0,)\varepsilon:E\to[0,\infty)80, and ε:E[0,)\varepsilon:E\to[0,\infty)81, then compute

ε:E[0,)\varepsilon:E\to[0,\infty)82

and output ε:E[0,)\varepsilon:E\to[0,\infty)83. The summary states that each entry of ε:E[0,)\varepsilon:E\to[0,\infty)84 and ε:E[0,)\varepsilon:E\to[0,\infty)85 has add–remove sensitivity ε:E[0,)\varepsilon:E\to[0,\infty)86 in ε:E[0,)\varepsilon:E\to[0,\infty)87, so the Laplace scale yields ε:E[0,)\varepsilon:E\to[0,\infty)88-AHDP.

The empirical section evaluates these mechanisms on two GPT-4o-generated synthetic datasets and one real dataset. For mean estimation, the synthetic data contain approximately ε:E[0,)\varepsilon:E\to[0,\infty)89 samples of ε:E[0,)\varepsilon:E\to[0,\infty)90 with correlation approximately ε:E[0,)\varepsilon:E\to[0,\infty)91; Sampling Mechanism ε:E[0,)\varepsilon:E\to[0,\infty)92 yields ε:E[0,)\varepsilon:E\to[0,\infty)93 as ε:E[0,)\varepsilon:E\to[0,\infty)94, whereas fixed-weight methods incur a constant bias. For frequency estimation, the data contain ε:E[0,)\varepsilon:E\to[0,\infty)95 GPT-generated tuples across ε:E[0,)\varepsilon:E\to[0,\infty)96 education levels and ε:E[0,)\varepsilon:E\to[0,\infty)97 ε:E[0,)\varepsilon:E\to[0,\infty)98-levels, with a ε:E[0,)\varepsilon:E\to[0,\infty)99 test rejecting independence; Sampling XXn\mathbf X\in\mathcal X^n00 outperforms but does not vanish as XXn\mathbf X\in\mathcal X^n01, while linear-query methods have nontrivial error. For linear regression on California Housing, with XXn\mathbf X\in\mathcal X^n02 training and XXn\mathbf X\in\mathcal X^n03 test samples and XXn\mathbf X\in\mathcal X^n04, weighted least squares with XXn\mathbf X\in\mathcal X^n05 has the lowest error among universal AHDP schemes and approaches the nonprivate baseline as XXn\mathbf X\in\mathcal X^n06 grows. The same summary reports computational costs of XXn\mathbf X\in\mathcal X^n07 for sums and histograms and XXn\mathbf X\in\mathcal X^n08 for regression.

6. Conceptual issues, relationships, and open directions

Several technical issues recur across the AHDP literature. One is the relation to personalized DP. The graph-based paper states that allowing XXn\mathbf X\in\mathcal X^n09 to be a function of neighboring datasets recovers an earlier definition of personalized DP as a special case (Torkamani et al., 2022). Another is the status of correlation between data and privacy demand. The correlation-aware paper makes this a central objection to standard HDP, asserting that frameworks which vary data while holding privacy labels fixed can be inadequate when XXn\mathbf X\in\mathcal X^n10 is correlated and private (Chaudhuri et al., 2 Sep 2025).

A common misconception is that heterogeneity merely means replacing a single XXn\mathbf X\in\mathcal X^n11 with a vector XXn\mathbf X\in\mathcal X^n12 and then applying standard intuitions unchanged. The papers collectively indicate that this is incomplete. In the graph formulation, heterogeneity propagates through paths and induced constraints, so compatibility and optimal extension are global graph problems rather than local budget substitutions. In mean estimation, the minimax rate is controlled by

XXn\mathbf X\in\mathcal X^n13

which encodes a prefix effect rather than simple averaging of budgets. In the correlation-aware formulation, the relevant protected object is the joint type XXn\mathbf X\in\mathcal X^n14, not merely the datum XXn\mathbf X\in\mathcal X^n15.

The papers also identify distinct bias–variance trade-offs. In graph-based AHDP, utility is framed as maximizing XXn\mathbf X\in\mathcal X^n16 at each vertex subject to heterogeneous DP, and the extension is pointwise optimal. In the mean-estimation setting, ADPM balances XXn\mathbf X\in\mathcal X^n17 against XXn\mathbf X\in\mathcal X^n18, yielding the saturation phenomenon in which larger-budget users may receive more privacy than requested without utility loss (Chaudhuri et al., 2023). In universal AHDP mechanisms for correlated data, linear-query methods trade bias from down-weighting private points against DP noise, while sampling trades subsampling bias against DP noise; the summary explicitly notes that the choice of XXn\mathbf X\in\mathcal X^n19 or XXn\mathbf X\in\mathcal X^n20 must be tuned (Chaudhuri et al., 2 Sep 2025).

Open problems are stated in task-specific terms. For heterogeneous mean estimation these include tighter heterogeneous-DP composition theorems, data-dependent tuning of weights when XXn\mathbf X\in\mathcal X^n21 are private, extensions to sub-Gaussian or heavy-tailed domains, multivariate AHDP, and practical implementations in streaming or federated settings (Chaudhuri et al., 2023). For correlation-aware AHDP, the listed directions include XXn\mathbf X\in\mathcal X^n22-AHDP or Rényi-AHDP variants via Gaussian mechanisms, task-specific mechanisms beyond linear models, membership-inference stress tests, characterization of minimal bias under realistic distributional assumptions, and optimal budget splitting between numerator and denominator in mean estimation (Chaudhuri et al., 2 Sep 2025).

Taken together, these works present AHDP as a technically heterogeneous research area organized around a shared principle: add/remove privacy guarantees can be individualized, but the correct formal object for that individualization depends on whether the heterogeneity is attached to graph edges, user coordinates, or joint data–privacy types. This suggests that the central unresolved question is not whether heterogeneous privacy is possible, but which add/remove formalism is appropriate for the statistical, structural, and adversarial assumptions of a given problem.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (3)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Add-remove Heterogeneous Differential Privacy (AHDP).