Ohm, P. (2009). Broken promises of privacy: Responding to the surprising failure of anonymization. SSRN.
In this article the author investigates the merits of anonymization. Anonymization is a now-standard practice where datasets are stripped of personally identifiable information (PII) such that they are agreed to be suitable for sharing, whether with a private party or the public. Ohm argues that anonymization is inadequate in the face of reidentification, a strategy where datasets are combined with external data and information to yield unique data fingerprints of individuals.
The author highlights key advantages to anonymization in a social, political, and legal framework. It allows a discrete action to take place before data can be shared. It allows this action to be enforced and legislated, and so consequently allows an easy division between responsible data-sharers and irresponsible parties. Anonymization balances interests between researchers, legislators, and the public.
Key to anonymization is the idea of PII. When PII are identified in a record (perhaps a name, ZIP, birth date, etc.), those values are either suppressed (removed) or generalized (for instance a full birth date is translated to just the year of birth) so that they are no longer deemed personally identifiable values.
Ohm presents several instances where datasets anonymized in this fashion were used to very accurately identify a single person or demonstrate the ease of this task: the 2006 AOL release of 20 million search queries that identified user 4417749 as Thelma Arnold, the release of Massachusetts state employee health records by the Group Insurance Commission that revealed the governor’s own health records (diagnoses and prescriptions included), and Netflix’s release of 100 million movie ratings records in an open contest to develop a better recommendation algorithm. In the latter case UT researchers were able to demonstrate how little outside data an investigator needed to identify individual’s profiles amongst these records.
The author goes on to suggest that “anonymization” and PII are not realistic terms. The former implies absolute success without varying degrees, and the latter fails to account for the accretion of non-PII fields to constitute a unique identity, which is the primary strategy of reidentification. The author also demonstrates how attempts to extend PII fields invariably decrease the value of data: the more anonymous data becomes, the more value it loses with researchers.
On this struggle the author contrasts the US’s approach to privacy-data legislation and the EU’s. The US largely functions sector-by-sector, addressing the health profession, scientific communities, state departments, etc., in turn. EU maintains a Data Protection Directive which seeks to regulate all data everywhere for PII elements. Both have serious drawbacks, since US legislation is liable to skip or under-regulate whole industries with the idea that its data cannot be used harmfully. The EU meanwhile faces a rising tide of reidentification instances that place more data in the PII camp, thus severely limiting sharing.
Last is a consideration of alternatives to legislating privacy in datasets. The author argues that comprehensive and sector-by-sector approaches need to be combined, and that a risk-assessment of data can guide the sector-by-sector guidelines. Some of the risk factors the author lists are: data handling techniques, whether data is moving to a private or public party, the sheer quantity of the data (since reidentification benefits considerably from large datasets), and motive (academic researchers have little motive to reidentify because of rules and professionalism).
We’ve mentioned privacy briefly in class, by I think this review of anonymization is relevant to digital curation since the field hinges on the idea of mass sharing of data that’s not necessarily tracked carefully. Fortunately the field consists of academic researchers and graduate students so it seems unlikely that in private-to-private data transactions much malevolent intent will occur. Anonymization will probably be used quite a bit in curation since it allows one to “share and forget.” So long as that’s the case it’s good to acknowledge that anonymized data has only been scrubbed to a certain degree, and in conjunction with other external datasets could achieve a lot more specificity than originally intended.