Papers
Topics
Authors
Recent
Search
2000 character limit reached

ODkAnon: Greedy k-Anonymization with H3

Updated 12 July 2026
  • ODkAnon is a novel greedy algorithm that generalizes spatial origin–destination matrices using the H3 hexagon hierarchy to achieve k-anonymity while preserving analytical utility.
  • It systematically aggregates sparse cell counts into homogeneous zones and employs a suppression strategy within a defined budget to maintain privacy boundaries.
  • Empirical evaluations show that ODkAnon balances participant- and population-level privacy, outperforming methods like OIGH and ATG-Soft in scalability and utility trade-offs.

ODkAnon most directly denotes a novel greedy algorithm for producing homogeneous kk-anonymous origin–destination (OD) matrices over a spatial hierarchy, introduced in the mobility-privacy literature as a method that generalizes sparse origin and destination cells upward until each released OD pair has at least kk trips (Armenante et al., 16 Sep 2025). In the supplied literature, the label also appears in broader or analogous senses: as a style of k-anonymous aggregate-by-design A/B testing, as a descriptor associated with the anjana Python library for tabular anonymization, and as an inverse threat model exemplified by OCEAN, which shows how weakly protected public data can be collated to reconstruct rich personal dossiers (Gershoff, 24 Jan 2025, Díaz et al., 2024, Gupta et al., 2013). This suggests a broader privacy-engineering theme: transforming sparse, linkable, or quasi-identifying data into generalized or aggregated forms that retain analytical value while reducing re-identification risk.

1. Nomenclature and scope

In its primary and most specific usage, ODkAnon is the anonymization method proposed for origin–destination matrices in mobility analysis. Its stated purpose is to turn a high-resolution mobility table into a kk-anonymous OD matrix while preserving as much analytical utility as possible, using Uber’s H3 hexagon hierarchy and a greedy spatial generalization strategy (Armenante et al., 16 Sep 2025).

The supplied literature, however, uses the label more broadly. In the description of "K-Anonymous A/B Testing", the approach is characterized as an “ODkAnon” style approach because it replaces individual-level telemetry with equivalence-class summaries and runs regression from aggregate sufficient statistics (Gershoff, 24 Jan 2025). In the description of "An Open Source Python Library for Anonymizing Sensitive Data", ODkAnon is presented under the name anjana, an open-source Python framework for local anonymization of sensitive tabular data through identifier removal, quasi-identifier generalization, suppression limits, and classic privacy models such as kk-anonymity, ℓ\ell-diversity, and tt-closeness (Díaz et al., 2024). In the description of "OCEAN: Open-source Collation of eGovernment data And Networks", the system is described as closely analogous to an ODkAnon-style threat, because it shows how minimal seed information can be expanded into a large dossier by collating open government databases and social-network traces (Gupta et al., 2013).

A separate acronym, ODK, refers to the Ontology Development Kit, a toolkit for building, maintaining, and standardising biomedical ontologies. It is distinct from ODkAnon, despite the superficial similarity of initials (Matentzoglu et al., 2022).

Referent Context Core mechanism
ODkAnon Mobility OD anonymization Greedy hierarchical generalization and suppression
“ODkAnon” style A/B testing Aggregate-by-design experimentation Equivalence-class summaries for OLS
ODkAnon/anjana Tabular anonymization Hierarchies, suppression, classic privacy models
ODkAnon-style threat Deanonymization analogue Cross-source collation of public data
ODK Ontology engineering Dockerized workflows and standardization

2. OD matrices, sparse mobility data, and the privacy target

The ODkAnon algorithm is motivated by a standard privacy problem in mobility analysis: even aggregated OD matrices can reveal sensitive movement patterns when cells are sparse, routes are unique, or flow counts are small. In this setting, a matrix cell summarizes trips from origin oo to destination dd, but low-count cells can still support re-identification or inference attacks (Armenante et al., 16 Sep 2025).

A central contribution of the ODkAnon paper is the distinction between participant-protecting anonymization and population-protecting anonymization. In the NetMob 2025 setting, each participant is not simply one record but a weighted proxy for a segment of the real population. The dataset includes a weight attribute, WEIGHT_INDVID, and the authors argue that this changes the privacy perspective fundamentally: a matrix can be kk-anonymous for survey participants yet fail to be kk-anonymous for the inferred population represented by those participants (Armenante et al., 16 Sep 2025).

The paper restates the usual intuition that a dataset is kk0-anonymous if every record cannot be distinguished among at least kk1 others, and in the OD setting this is interpreted as requiring at least kk2 trips from the same origin to the same destination. Once weights are introduced, the same trip may count many times according to representativeness, so privacy must be evaluated over the expanded population rather than only over the sampled records (Armenante et al., 16 Sep 2025).

The matrix setting is further complicated by socio-demographic segmentation. The dataset supports OD matrices stratified by sex, age, and socio-professional category. This matters because anonymization difficulty is not uniform across groups: some segments are sparser and more distinctive, and thus require more aggressive generalization or suppression. The reported results show that these differences are substantive rather than cosmetic; protecting men, women, age bands, or socio-professional categories can produce markedly different anonymization patterns and utility losses (Armenante et al., 16 Sep 2025).

3. Algorithmic structure of ODkAnon

ODkAnon is presented as a greedy algorithm for kk3-anonymous OD matrices built on the H3 hierarchy. Its conceptual strategy is to start from fine-grained H3 cells, identify sparse groups, and repeatedly generalize them to parent cells until all remaining OD cells satisfy the threshold kk4, or else suppress rows that cannot be salvaged within a specified budget (Armenante et al., 16 Sep 2025).

The main inputs are the OD matrix, the anonymity threshold kk5, the H3 hierarchy, a maximum generalization depth kk6, and a suppression budget kk7. The output is a filtered and generalized OD matrix whose surviving cells satisfy kk8-anonymity, together with suppressed rows or trips as needed (Armenante et al., 16 Sep 2025).

Operationally, the method:

  1. builds hierarchical trees for origins and destinations using H3;
  2. initializes a sparse OD matrix, explicitly using CSR/CSC-style sparse handling;
  3. precomputes sibling groups of cells sharing a parent;
  4. iteratively generalizes along one axis;
  5. merges sibling groups into parents and updates the sparse matrix;
  6. stops when all cells meet kk9, or no valid generalization remains;
  7. suppresses remaining problematic rows if necessary and within budget (Armenante et al., 16 Sep 2025).

A notable design choice is the enforcement of homogeneous areas. The generalized origin zones do not depend on destination, and the generalized destination zones do not depend on origin. This contrasts with more flexible but less interpretable schemes in which the representation of a zone can vary by OD pair. The paper explicitly argues that homogeneous zones are easier to interpret and use in downstream mobility analysis (Armenante et al., 16 Sep 2025).

The balancing heuristic is also specific. ODkAnon tracks the ratio between the number of origins and destinations. If that ratio deviates by more than kk0 from the initial value, generalization is forced on the dominant axis; otherwise the algorithm alternates between axes. Within the selected axis, it chooses the sibling group with the lowest aggregated count as the next group to generalize. This is a greedy “fix the smallest violation first” rule aimed at limiting unnecessary information loss (Armenante et al., 16 Sep 2025).

The method includes a separate suppression algorithm because some OD pairs remain too sparse even after hierarchical lifting. For each OD pair, the algorithm explores parent hexagons from fine to coarser levels up to kk1, checks whether the aggregated count reaches kk2, marks rows that never reach kk3 as problematic, and suppresses them if the number of problematic rows is within the suppression budget kk4. If it is not, the algorithm suppresses only the rows with the lowest counts (Armenante et al., 16 Sep 2025).

The tree machinery is defined over H3 cells by extracting unique hexagons, determining a root resolution via the minimal optimal resolution kk5, building parent–child relationships up to kk6, and propagating trip counts upward so that each node stores the total trips in its subtree. Separate trees are maintained for starts and ends, denoted tree_start and tree_end (Armenante et al., 16 Sep 2025).

4. Evaluation criteria and empirical behavior

The ODkAnon paper evaluates privacy protection together with utility degradation, and compares ODkAnon against ATG-Soft, OIGH, and Mondrian (Armenante et al., 16 Sep 2025).

The privacy evaluation explicitly checks what happens when anonymization is optimized for one target and then evaluated for the other. The main conclusion is that protecting participants does not guarantee protecting the population, and conversely that protecting the inferred population does not reduce to participant-level anonymity (Armenante et al., 16 Sep 2025).

Utility is measured with four metrics:

kk7

kk8

kk9

kk0

Here kk1 is the Discernability Metric, kk2 the Normalized Average Equivalence Class Size, kk3 the Mean Generalization Error, and kk4 the Reconstruction Loss (Armenante et al., 16 Sep 2025).

The comparison is constrained by a two-hour runtime limit per run. Under that limit, the paper reports that OIGH is faster than ODkAnon but often loses substantially more utility because it cannot use suppression and therefore tends to over-generalize sparse cells. ATG-Soft is more flexible and can produce non-homogeneous generalization, but it often has poor scalability and may fail to finish within the time limit. Mondrian can perform well on some general utility metrics such as kk5 and kk6, but because it partitions space into bounding hypercubes or rectangles rather than using the H3 hierarchy, kk7 and kk8 are not defined for it, and it produces non-homogeneous and overlapping regions (Armenante et al., 16 Sep 2025).

The reported findings position ODkAnon as a compromise between feasibility and utility. It usually offers better utility than ATG-Soft and OIGH, especially under the homogeneous-zone constraint, while remaining computationally feasible in scenarios where ATG-Soft struggles. On the whole dataset, the paper reports that ODkAnon produced 29 zones in both origins and destinations when protecting participants with kk9, whereas protecting the population led to a different structure, for example 35 origin zones and 29 destination zones, reflecting the effect of weighted representativeness (Armenante et al., 16 Sep 2025).

Outside mobility, the supplied literature uses ODkAnon as a broader pattern of aggregate-first anonymization. In A/B testing, the core idea is that many analyses normally performed on microdata can instead be done on equivalence classes defined by quasi-identifiers such as treatment assignment and categorical covariates. Each class stores at least the quasi-identifier values, the count in the class, and the sum of the outcome, with optional totals such as ℓ\ell0 when variance estimation requires them. Because OLS depends on ℓ\ell1 and ℓ\ell2, regression can be reconstructed exactly from those aggregates for the models considered (Gershoff, 24 Jan 2025).

That paper emphasizes that k-anonymity is not a formal privacy guarantee like differential privacy, but treats it as a simple and auditable data-minimization measure. The approach supports use cases such as partial F-tests for detecting interactions or heterogeneous treatment effects, and regression adjustment using a CUPED-like ANCOVA formulation. Its stated advantages include privacy by design, lower storage and compute costs, simpler governance, and exact or near-exact OLS inference from aggregated data (Gershoff, 24 Jan 2025).

In tabular anonymization, the supplied description associates ODkAnon with the anjana Python library. That framework assumes columns are partitioned into identifiers, quasi-identifiers, sensitive attributes, and insensitive attributes. Identifiers are removed or replaced with *, quasi-identifiers are generalized through user-defined hierarchies, and the library enforces one of nine classic privacy models: k-anonymity, (ℓ\ell3,k)-anonymity, ℓ\ell4-diversity, entropy ℓ\ell5-diversity, recursive (c,ℓ\ell6)-diversity, t-closeness, ℓ\ell7-disclosure privacy, basic ℓ\ell8-likeness, and enhanced ℓ\ell9-likeness. The library is local, open source, Python-based, and designed to fit ML/DL workflows (Díaz et al., 2024).

An inverse use of the same conceptual space appears in OCEAN, which is not an anonymization system but a deanonymization / privacy-leak / re-identification system. It takes a small amount of seed information such as a name or name plus location, queries public sources including the Delhi Driving Licence database, Delhi Voter ID / electoral-roll database, PAN card status database, MTNL phone directory, and public APIs from Facebook, Twitter, Foursquare, LinkedIn, and Google Plus, and returns a much larger dossier containing attributes such as name, age, address, date of birth, parents’ names, voter ID, driving licence number, and PAN. The authors describe the same-source chaining of records as enabling “horizontal-depth” expansion (Gupta et al., 2013).

This contrast is informative. ODkAnon in the anonymization sense attempts to make sparse observations indistinguishable through aggregation, whereas OCEAN demonstrates how publicly exposed and linkable records can be chained together to defeat effective anonymity. A plausible implication is that the ODkAnon family and OCEAN occupy opposite ends of the same privacy pipeline: one minimizes identifiability in release, the other exploits residual identifiability in exposed data (Gupta et al., 2013).

6. Limitations, assumptions, and broader significance

The primary ODkAnon method inherits several limitations from the OD-matrix setting. It is designed for sparse, hierarchically organized mobility data, and its behavior depends on the available spatial hierarchy, the suppression budget, and the chosen anonymity target. The paper is explicit that participant-level and population-level protection are not interchangeable, and that socio-demographic segmentation can make anonymization difficulty highly uneven across subpopulations (Armenante et al., 16 Sep 2025).

The aggregate-by-design A/B testing approach is likewise bounded by what can be recovered from class-level sufficient statistics. It is best suited to analyses for which tt0, tt1, and the necessary sums of squares can be computed from equivalence classes. More detailed models may require finer-grained storage, weakening tt2-anonymity. The author also stresses that k-anonymity is a practical minimization strategy rather than a formal privacy guarantee (Gershoff, 24 Jan 2025).

For tabular anonymization, the anjana description identifies several practical constraints: version 1.0.0 supports one sensitive attribute, it works on tabular data, and it requires users to provide appropriate generalization hierarchies. Achievable privacy depends strongly on hierarchy design, and some privacy targets may fail under a given hierarchy configuration. In the reported example, some demanding methods, especially entropy tt3-diversity and recursive (c,tt4)-diversity, could not be satisfied with the particular hierarchies used (Díaz et al., 2024).

The OCEAN threat model highlights the operational consequences of failing to apply such controls. The paper states that exposed PII can be used to create fake documents, open fake bank accounts, procure phone connections or credit cards, impersonate individuals, and register them on vulnerable portals to retrieve even more sensitive information. It applies Microsoft’s DREAD model and assigns a total risk score of 13, describing the risk as high. The proposed defenses include authorization mechanisms such as usernames and passwords, CAPTCHA to slow automated scraping, stronger privacy laws, and restrictions on direct public access to uniquely identifying fields such as voter ID, driving licence number, PAN, address, and phone number (Gupta et al., 2013).

Taken together, these works place ODkAnon within a technically coherent, though terminologically heterogeneous, area of privacy engineering. In its strict sense, ODkAnon is a greedy H3-based method for publishing homogeneous tt5-anonymous OD matrices under both participant and population threat models (Armenante et al., 16 Sep 2025). In a broader editorial sense, the term denotes a family resemblance among methods that rely on equivalence classes, hierarchical generalization, suppression, and aggregate sufficient statistics to reduce identifiability while retaining utility (Gershoff, 24 Jan 2025, Díaz et al., 2024). The opposing example of OCEAN clarifies why these methods matter: when data are left open, linkable, and weakly controlled, sparse identifiers can be amplified into high-risk personal dossiers (Gupta et al., 2013).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ODkAnon.