Skip to content
Vol. 1 · Ed. 2026
CyberGlossary
Entry № 306

Data Anonymization

Reviewed byCybersecurity entrepreneur & security researcher

What is Data Anonymization?

Data AnonymizationIrreversibly transforming personal data so that no individual can be identified, directly or indirectly, even when combined with other available information.


Data anonymization irreversibly transforms personal data so that individuals can no longer be singled out, linked, or inferred — even by combining the release with outside information. Core techniques are suppression, generalization (e.g. coarsening a birth date to a year), aggregation, perturbation/noise, and randomization, usually evaluated against a formal model: k-anonymity, l-diversity, t-closeness, or differential privacy, which adds calibrated noise under a provable privacy budget (ε).

The bar is high because of a long record of re-identification. Latanya Sweeney famously estimated that 87% of Americans are unique on the combination of 5-digit ZIP, date of birth, and sex, and used public voter rolls to re-identify Massachusetts Governor William Weld's hospital records in "anonymized" state data — the result that motivated k-anonymity. In 2008 Narayanan and Shmatikov de-anonymized the Netflix Prize dataset (ratings of ~500,000 subscribers) by correlating sparse ratings with public IMDb profiles. The 2006 AOL search-log release was traced back to named individuals within days.

Under the GDPR, truly anonymized data falls outside the Regulation (Recital 26), but the Article 29 Working Party's Opinion 05/2014 and regulators such as the EDPB and CNIL require a documented re-identification risk assessment covering singling-out, linkability, and inference using means "reasonably likely to be used." Common pitfalls: relying on hashing or simple pseudonyms, releasing high-dimensional micro-data, or confusing pseudonymization with anonymization.

flowchart TD
  R["Raw personal data"] --> D["Remove direct identifiers"]
  D --> Q["Generalize / suppress quasi-identifiers"]
  Q --> M["Apply privacy model (k-anonymity / differential privacy)"]
  M --> T{"Re-identification risk acceptable?"}
  T -->|"No"| Q
  T -->|"Yes"| P["Publish / share anonymized dataset"]

● Examples

  1. 01

    Publishing hospital readmission statistics aggregated by region and quarter, with cells below five suppressed.

  2. 02

    Releasing a public mobility dataset where trajectories are generalized to neighborhood-week granularity.

● Frequently asked questions

What is Data Anonymization?

Irreversibly transforming personal data so that no individual can be identified, directly or indirectly, even when combined with other available information. It belongs to the Privacy & Data Protection category of cybersecurity.

What does Data Anonymization mean?

Irreversibly transforming personal data so that no individual can be identified, directly or indirectly, even when combined with other available information.

How do you defend against Data Anonymization?

Defences for Data Anonymization typically combine technical controls and operational practices, as detailed in the full definition above.

What are other names for Data Anonymization?

Common alternative names include: Anonymization, De-identification (strong sense).

● Related terms

● See also