---
title: 'ODkAnon: Greedy k-Anonymization with H3'
url: https://www.emergentmind.com/topics/odkanon
type: topic
---

# ODkAnon: Greedy k-Anonymization with H3

ODkAnon most directly denotes a **novel greedy algorithm** for producing **homogeneous \(k\)-anonymous origin–destination (OD) matrices** over a spatial hierarchy, introduced in the mobility-privacy literature as a method that generalizes sparse origin and destination cells upward until each released OD pair has at least \(k\) trips [2509.12950]. In the supplied literature, the label also appears in broader or analogous senses: as a style of **k-anonymous aggregate-by-design** A/B testing, as a descriptor associated with the **anjana** Python library for tabular anonymization, and as an inverse threat model exemplified by **OCEAN**, which shows how weakly protected public data can be collated to reconstruct rich personal dossiers [2501.14329] [2408.10766] [1312.2784]. This suggests a broader privacy-engineering theme: transforming sparse, linkable, or quasi-identifying data into generalized or aggregated forms that retain analytical value while reducing re-identification risk.

## 1. Nomenclature and scope

In its primary and most specific usage, ODkAnon is the anonymization method proposed for **origin–destination matrices** in mobility analysis. Its stated purpose is to turn a high-resolution mobility table into a **\(k\)-anonymous OD matrix** while preserving as much analytical utility as possible, using Uber’s **H3 hexagon hierarchy** and a greedy spatial generalization strategy [2509.12950].

The supplied literature, however, uses the label more broadly. In the description of **"K-Anonymous A/B Testing"**, the approach is characterized as an **“ODkAnon” style approach** because it replaces individual-level telemetry with **equivalence-class summaries** and runs regression from aggregate sufficient statistics [2501.14329]. In the description of **"An Open Source Python Library for Anonymizing Sensitive Data"**, ODkAnon is presented under the name **anjana**, an open-source Python framework for local anonymization of sensitive tabular data through identifier removal, quasi-identifier generalization, suppression limits, and classic privacy models such as \(k\)-anonymity, \(\ell\)-diversity, and \(t\)-closeness [2408.10766]. In the description of **"OCEAN: Open-source Collation of eGovernment data And Networks"**, the system is described as closely analogous to an **ODkAnon-style threat**, because it shows how minimal seed information can be expanded into a large dossier by collating open government databases and social-network traces [1312.2784].

A separate acronym, **ODK**, refers to the **Ontology Development Kit**, a toolkit for building, maintaining, and standardising biomedical ontologies. It is distinct from ODkAnon, despite the superficial similarity of initials [2207.02056].

| Referent | Context | Core mechanism |
|---|---|---|
| ODkAnon | Mobility OD anonymization | Greedy hierarchical generalization and suppression |
| “ODkAnon” style A/B testing | Aggregate-by-design experimentation | Equivalence-class summaries for OLS |
| ODkAnon/anjana | Tabular anonymization | Hierarchies, suppression, classic privacy models |
| ODkAnon-style threat | Deanonymization analogue | Cross-source collation of public data |
| ODK | Ontology engineering | Dockerized workflows and standardization |

## 2. OD matrices, sparse mobility data, and the privacy target

The ODkAnon algorithm is motivated by a standard privacy problem in mobility analysis: **even aggregated OD matrices can reveal sensitive movement patterns** when cells are sparse, routes are unique, or flow counts are small. In this setting, a matrix cell summarizes trips from origin \(o\) to destination \(d\), but low-count cells can still support re-identification or inference attacks [2509.12950].

A central contribution of the ODkAnon paper is the distinction between **participant-protecting anonymization** and **population-protecting anonymization**. In the NetMob 2025 setting, each participant is not simply one record but a **weighted proxy** for a segment of the real population. The dataset includes a weight attribute, `WEIGHT_INDVID`, and the authors argue that this changes the privacy perspective fundamentally: a matrix can be \(k\)-anonymous for survey participants yet fail to be \(k\)-anonymous for the **inferred population** represented by those participants [2509.12950].

The paper restates the usual intuition that a dataset is \(k\)-anonymous if every record cannot be distinguished among at least \(k-1\) others, and in the OD setting this is interpreted as requiring **at least \(k\) trips from the same origin to the same destination**. Once weights are introduced, the same trip may count many times according to representativeness, so privacy must be evaluated over the expanded population rather than only over the sampled records [2509.12950].

The matrix setting is further complicated by **socio-demographic segmentation**. The dataset supports OD matrices stratified by **sex**, **age**, and **socio-professional category**. This matters because anonymization difficulty is not uniform across groups: some segments are sparser and more distinctive, and thus require more aggressive generalization or suppression. The reported results show that these differences are substantive rather than cosmetic; protecting men, women, age bands, or socio-professional categories can produce markedly different anonymization patterns and utility losses [2509.12950].

## 3. Algorithmic structure of ODkAnon

ODkAnon is presented as a **greedy algorithm** for \(k\)-anonymous OD matrices built on the **H3 hierarchy**. Its conceptual strategy is to start from fine-grained H3 cells, identify sparse groups, and repeatedly generalize them to parent cells until all remaining OD cells satisfy the threshold \(k\), or else suppress rows that cannot be salvaged within a specified budget [2509.12950].

The main inputs are the **OD matrix**, the anonymity threshold \(k\), the **H3 hierarchy**, a **maximum generalization depth \(L\)**, and a **suppression budget \(\beta\)**. The output is a **filtered and generalized OD matrix** whose surviving cells satisfy \(k\)-anonymity, together with suppressed rows or trips as needed [2509.12950].

Operationally, the method:

1. builds **hierarchical trees** for origins and destinations using H3;
2. initializes a **sparse OD matrix**, explicitly using CSR/CSC-style sparse handling;
3. precomputes **sibling groups** of cells sharing a parent;
4. iteratively generalizes along one axis;
5. merges sibling groups into parents and updates the sparse matrix;
6. stops when all cells meet \(k\), or no valid generalization remains;
7. suppresses remaining problematic rows if necessary and within budget [2509.12950].

A notable design choice is the enforcement of **homogeneous areas**. The generalized origin zones do **not** depend on destination, and the generalized destination zones do **not** depend on origin. This contrasts with more flexible but less interpretable schemes in which the representation of a zone can vary by OD pair. The paper explicitly argues that homogeneous zones are easier to interpret and use in downstream mobility analysis [2509.12950].

The **balancing heuristic** is also specific. ODkAnon tracks the ratio between the number of origins and destinations. If that ratio deviates by more than \(\pm 3\%\) from the initial value, generalization is forced on the dominant axis; otherwise the algorithm alternates between axes. Within the selected axis, it chooses the **sibling group with the lowest aggregated count** as the next group to generalize. This is a greedy “fix the smallest violation first” rule aimed at limiting unnecessary information loss [2509.12950].

The method includes a separate **suppression algorithm** because some OD pairs remain too sparse even after hierarchical lifting. For each OD pair, the algorithm explores parent hexagons from fine to coarser levels up to \(L\), checks whether the aggregated count reaches \(k\), marks rows that never reach \(k\) as problematic, and suppresses them if the number of problematic rows is within the suppression budget \(\beta\). If it is not, the algorithm suppresses only the rows with the **lowest counts** [2509.12950].

The tree machinery is defined over H3 cells by extracting unique hexagons, determining a root resolution via the **minimal optimal resolution \(R_{\min}\)**, building parent–child relationships up to \(R_{\text{target}}\), and propagating trip counts upward so that each node stores the total trips in its subtree. Separate trees are maintained for starts and ends, denoted `tree_start` and `tree_end` [2509.12950].

## 4. Evaluation criteria and empirical behavior

The ODkAnon paper evaluates privacy protection together with utility degradation, and compares ODkAnon against **ATG-Soft**, **OIGH**, and **Mondrian** [2509.12950].

The privacy evaluation explicitly checks what happens when anonymization is optimized for one target and then evaluated for the other. The main conclusion is that **protecting participants does not guarantee protecting the population**, and conversely that protecting the inferred population does not reduce to participant-level anonymity [2509.12950].

Utility is measured with four metrics:

\[
C_{DM} = \sum_{Eq \text{ s.t. } |Eq| \geq k} |Eq|^2 + \sum_{Eq \text{ s.t. } |Eq| < k}|D||Eq|
\]

\[
C_{AVG} = \dfrac{\left(\dfrac{|D^+|}{total\_equiv\_classes}\right)}{k}
\]

\[
\bar{G} = \frac{1}{|D^+|}\sum_{Eq_{o\xrightarrow{}d} \text{ s.t. } \left|Eq_{o\xrightarrow{}d}\right| \geq k}(|o| + |d|)\left|Eq_{o\xrightarrow{}d}\right|
\]

\[
E = \frac{1}{|D|} \sum_{o,d \in leaves(T)} \left|\tilde{D}_{o \xrightarrow{} d} - D_{o \xrightarrow{} d}\right|
\]

Here \(C_{DM}\) is the **Discernability Metric**, \(C_{AVG}\) the **Normalized Average Equivalence Class Size**, \(\bar{G}\) the **Mean Generalization Error**, and \(E\) the **Reconstruction Loss** [2509.12950].

The comparison is constrained by a **two-hour** runtime limit per run. Under that limit, the paper reports that **OIGH is faster** than ODkAnon but often loses substantially more utility because it cannot use suppression and therefore tends to over-generalize sparse cells. **ATG-Soft** is more flexible and can produce non-homogeneous generalization, but it often has poor scalability and may fail to finish within the time limit. **Mondrian** can perform well on some general utility metrics such as \(C_{DM}\) and \(C_{AVG}\), but because it partitions space into bounding hypercubes or rectangles rather than using the H3 hierarchy, \(\bar{G}\) and \(E\) are not defined for it, and it produces non-homogeneous and overlapping regions [2509.12950].

The reported findings position ODkAnon as a compromise between feasibility and utility. It **usually offers better utility than ATG-Soft and OIGH**, especially under the homogeneous-zone constraint, while remaining computationally feasible in scenarios where ATG-Soft struggles. On the whole dataset, the paper reports that ODkAnon produced **29 zones** in both origins and destinations when protecting participants with \(k=10\), whereas protecting the population led to a different structure, for example **35 origin zones and 29 destination zones**, reflecting the effect of weighted representativeness [2509.12950].

## 5. Related privacy-engineering uses of the ODkAnon pattern

Outside mobility, the supplied literature uses ODkAnon as a broader pattern of **aggregate-first anonymization**. In **A/B testing**, the core idea is that many analyses normally performed on microdata can instead be done on **equivalence classes** defined by quasi-identifiers such as treatment assignment and categorical covariates. Each class stores at least the quasi-identifier values, the **count** in the class, and the **sum of the outcome**, with optional totals such as \(\sum y_i^2\) when variance estimation requires them. Because OLS depends on \(X'X\) and \(X'y\), regression can be reconstructed exactly from those aggregates for the models considered [2501.14329].

That paper emphasizes that **k-anonymity is not a formal privacy guarantee like differential privacy**, but treats it as a simple and auditable data-minimization measure. The approach supports use cases such as **partial F-tests** for detecting interactions or heterogeneous treatment effects, and **regression adjustment** using a CUPED-like ANCOVA formulation. Its stated advantages include privacy by design, lower storage and compute costs, simpler governance, and exact or near-exact OLS inference from aggregated data [2501.14329].

In **tabular anonymization**, the supplied description associates ODkAnon with the **anjana** Python library. That framework assumes columns are partitioned into **identifiers**, **quasi-identifiers**, **sensitive attributes**, and **insensitive attributes**. Identifiers are removed or replaced with `*`, quasi-identifiers are generalized through user-defined hierarchies, and the library enforces one of nine classic privacy models: **k-anonymity**, **(\(\alpha\),k)-anonymity**, **\(\ell\)-diversity**, **entropy \(\ell\)-diversity**, **recursive (c,\(\ell\))-diversity**, **t-closeness**, **\(\delta\)-disclosure privacy**, **basic \(\beta\)-likeness**, and **enhanced \(\beta\)-likeness**. The library is local, open source, Python-based, and designed to fit ML/DL workflows [2408.10766].

An inverse use of the same conceptual space appears in **OCEAN**, which is not an anonymization system but a **deanonymization / privacy-leak / re-identification system**. It takes a small amount of seed information such as a **name** or **name plus location**, queries public sources including the **Delhi Driving Licence database**, **Delhi Voter ID / electoral-roll database**, **PAN card status database**, **MTNL phone directory**, and public APIs from **Facebook**, **Twitter**, **Foursquare**, **LinkedIn**, and **Google Plus**, and returns a much larger dossier containing attributes such as name, age, address, date of birth, parents’ names, voter ID, driving licence number, and PAN. The authors describe the same-source chaining of records as enabling **“horizontal-depth” expansion** [1312.2784].

This contrast is informative. ODkAnon in the anonymization sense attempts to make sparse observations indistinguishable through aggregation, whereas OCEAN demonstrates how publicly exposed and linkable records can be chained together to defeat effective anonymity. A plausible implication is that the ODkAnon family and OCEAN occupy opposite ends of the same privacy pipeline: one minimizes identifiability in release, the other exploits residual identifiability in exposed data [1312.2784].

## 6. Limitations, assumptions, and broader significance

The primary ODkAnon method inherits several limitations from the OD-matrix setting. It is designed for **sparse**, **hierarchically organized** mobility data, and its behavior depends on the available spatial hierarchy, the suppression budget, and the chosen anonymity target. The paper is explicit that participant-level and population-level protection are **not interchangeable**, and that socio-demographic segmentation can make anonymization difficulty highly uneven across subpopulations [2509.12950].

The aggregate-by-design A/B testing approach is likewise bounded by what can be recovered from class-level sufficient statistics. It is best suited to analyses for which \(X'X\), \(X'y\), and the necessary sums of squares can be computed from equivalence classes. More detailed models may require finer-grained storage, weakening \(k\)-anonymity. The author also stresses that k-anonymity is a practical minimization strategy rather than a formal privacy guarantee [2501.14329].

For tabular anonymization, the anjana description identifies several practical constraints: **version 1.0.0 supports one sensitive attribute**, it works on **tabular data**, and it requires users to provide appropriate **generalization hierarchies**. Achievable privacy depends strongly on hierarchy design, and some privacy targets may fail under a given hierarchy configuration. In the reported example, some demanding methods, especially **entropy \(\ell\)-diversity** and **recursive (c,\(\ell\))-diversity**, could not be satisfied with the particular hierarchies used [2408.10766].

The OCEAN threat model highlights the operational consequences of failing to apply such controls. The paper states that exposed PII can be used to create fake documents, open fake bank accounts, procure phone connections or credit cards, impersonate individuals, and register them on vulnerable portals to retrieve even more sensitive information. It applies Microsoft’s **DREAD** model and assigns a **total risk score of 13**, describing the risk as **high**. The proposed defenses include authorization mechanisms such as usernames and passwords, **CAPTCHA** to slow automated scraping, stronger privacy laws, and restrictions on direct public access to uniquely identifying fields such as voter ID, driving licence number, PAN, address, and phone number [1312.2784].

Taken together, these works place ODkAnon within a technically coherent, though terminologically heterogeneous, area of privacy engineering. In its strict sense, ODkAnon is a greedy H3-based method for publishing **homogeneous \(k\)-anonymous OD matrices** under both participant and population threat models [2509.12950]. In a broader editorial sense, the term denotes a family resemblance among methods that rely on **equivalence classes**, **hierarchical generalization**, **suppression**, and **aggregate sufficient statistics** to reduce identifiability while retaining utility [2501.14329] [2408.10766]. The opposing example of OCEAN clarifies why these methods matter: when data are left open, linkable, and weakly controlled, sparse identifiers can be amplified into high-risk personal dossiers [1312.2784].

Source: https://www.emergentmind.com/topics/odkanon