Issue 02 · Re-identification · 16 June 2026

De-identified does not mean what it used to

Every privacy policy lets a company pass on what it holds once your name is off it. The law allowed that because putting a name back took an expert, days and money. It now takes an AI model and a few seconds, and years of messages and profiles went out under the old assumption.

What the words mean

Four words do most of the work in a privacy policy, and they do not mean the same thing.

  • Anonymised means nobody can work out who you are from the data, by any means anyone is reasonably likely to use. Data that is truly anonymous sits outside data protection law altogether.
  • Pseudonymised means your name and the other direct identifiers have been swapped for a code, and the key that links the code back to you is kept somewhere else. Under UK and EU law it is still your personal data, and the rules still apply to it.
  • De-identified is the word most policies use, in the UK as much as the US, and some say the same thing without it: the data they keep cannot be tied back to you on its own. A record that cannot identify you on its own can still be matched with one that can, and the clause says nothing about that. Under the US health privacy rule a record counts as de-identified once eighteen listed items have been removed, or an expert has judged the risk of matching it back to be very small. California's law adds a promise: the business says it will not try to put the name back.
  • Aggregated means the data has been rolled up into totals about a group, such as how many customers in one city bought a thing this month.

In practice the words get used loosely, and often as if they were the same. "Anonymised" in a policy, or in a reply to a deletion request, most often means de-identified: the name and the email are gone, everything else is intact, and the record can still be matched back to you. That used to take real effort. It now takes an AI model and a few seconds. That is what this article is about.

What de-identified leaves behind

Personally identifiable information the law makes them protect this
Name Jane Smith
Address 184 Berry St, Brooklyn
Photos 3 on file
ID document driving licence
Everything else not personal on its own
ZIP code 11211
Date of birth 14 March 1999
Sex F
Education Bachelor's, Communications
Employer a media agency, Manhattan
Weekday area Williamsburg, Midtown
Weekend route Prospect Park loop
Evenings active 6PM to 2AM
Purchases skincare, fashion
Films rated 212
Phone iPhone 14
Browser Chrome on a Mac

Your record, in two parts.
Four of them name you.

The four on the left name you on their own, as an email address, a phone number or a passport would. The law makes companies protect these, and when you ask to see or delete your data, they are what the company looks for. Each piece on the right is shared with thousands of other people. On its own it does not identify you, and the de-identified clause treats it that way.

1 / 5

Step through it. The joins are the kind the studies below describe.

The paragraph that does the work

Somewhere in every privacy policy there is a paragraph that lets the company pass on what it holds about you, to partners, advertisers or buyers, as long as the data is "de-identified", "aggregated", "pseudonymised" or "anonymised" first. The words sound final. They describe one step: your name and email come off, and everything else stays on, including the messages, the profile, the timestamps, the postcode, the device and your search history.

The trade in personal data runs on that paragraph. Brokers buy and sell under it. "We share with partners" is the polite name for a sale, and the loose terms are what let the sale happen.

Data about you is either useful or anonymous; it cannot be both. A buyer pays for data because it describes someone, and the more detail it carries, the easier it is to tell who that someone is.

Why the law let the words stay loose

The law's definition of anonymous is written in terms of effort. The GDPR says that to decide whether a person could be identified from a set of data, you weigh "the costs of and the amount of time required for identification, taking into consideration the available technology at the time of the processing and technological developments". If nobody could reasonably afford to put your name back, the data counted as anonymous and the rules stopped applying to it.

The American versions rest on the same assumption. The US health privacy rule treats a record as de-identified once eighteen listed things are removed, a checklist with no test of whether the record can be matched back. California's definition is a promise and a contract: the business says it will not try to re-identify you, and makes whoever it passes the data to say the same.

Once the eighteen things are removed, a record is no longer protected health information and the rule stops applying to it: it can be shared, sold or published, and the rule asks nothing of whoever receives it. The company applies the list to itself. The one condition beyond the list is that it has no "actual knowledge" that what is left could still identify someone, and nobody outside the company checks the work. The regulator comes in on a complaint or a breach report.

California's promise is made by the business, to nobody in particular. The law asks it to commit in public that it will not try to put the names back, which is usually a line in its privacy policy, and to write the same commitment into its contract with whoever buys the data. Nobody has to show that the data cannot be matched back. You do not see the contract, and if the buyer breaks it, that is between the buyer, the seller and a regulator that would first have to find out.

In the UK and the EU the test is different and lands in the same place. Data is personal if whoever holds it could reasonably identify the person, and the holder answers that question about itself. The EU's top court said in 2025 that pseudonymised data "must not be regarded as constituting, in all cases and for every person, personal data": the same file can be personal data in one company's hands and not in another's, depending on what each can reasonably do with it. Nothing requires a company to announce that it has run a match. That happens inside its own systems, and the only trace on the outside is an advert or a decision that fits a little too well. The label a company gives a file also decides whether the rules on sharing it, and on sending it out of the country, apply at all, and the company chooses the label.

In each of the three, the safeguard is something the holder says about itself: that it struck the list, that it will not try, that it could not reasonably manage it.

The three definitions were written for a world where putting a name back was a project, which was a fair assumption at the time.

What it took, when someone bothered

In 2006 AOL published the search histories of hundreds of thousands of its users with the names replaced by numbers. The New York Times took user 4417749's searches, from "numb fingers" to "60 single men", and found a 62-year-old widow in Georgia. It took a newspaper, and it found one person.

In 2008 two researchers showed that "anonymised" Netflix viewing histories could be matched against public ratings on a film site to identify subscribers, which took a research team months of work on two neat tables of ratings, one from Netflix and one public.

In 2012 a team of computer scientists tried to identify anonymous writers by style alone, across 100,000 authors. Their first guess was right in about a fifth of cases.

In 2013 the Sunday Times asked a specialist to compare a debut crime novel against known authors, then had a second academic check the work, before it put to J. K. Rowling that she had written it.

Each of the four was possible, and each took people, a budget and time. Paying that to find one ordinary person was never worth it.

Every one of the four was in a newspaper or a published paper. The recital the law rests on asks a company to weigh "technological developments" alongside the tools of the day, and these were the developments. A company that passed on messages or profiles under the loose terms after 2012 did so with the method on the record and only the price in its favour.

What it takes now

In 2023 researchers at ETH Zurich gave language models ordinary forum posts and asked where the writers lived, what they earned and whether they were men or women. The models were right first time up to 85 percent of the time, "at a fraction of the cost (100x) and time (240x) required by humans", and stripping identifiers out of the text did little to stop them.

In early 2026 a team of security researchers built an AI agent that works on raw text across any platform, with no neat tables at all. It matched pseudonymous accounts to real people at up to 68 percent recall at 90 percent precision, which means it found about two thirds of the matches and was right nine times in ten, where the best older method managed close to none. One test split a single Reddit user's history in two and let the agent match the halves, and it did, since a person's own writing links their accounts to each other. The authors' conclusion is that "the practical obscurity protecting pseudonymous users online no longer holds".

Another 2026 paper opens with the assumption itself: anonymisation is trusted "because re-identification has historically required specialized expertise, tailored algorithms, and manual corroboration". Its AI agents reconstructed 79 percent of identities in the hardest version of the Netflix test, against 56 percent for the classical method, and matched redacted chat transcripts to specific people by checking the details against the public web.

In December 2025 Anthropic published 1,250 anonymised interviews with professionals for research. Within weeks a researcher using an AI model with web search had put names to six of the twenty-four scientists in the subset they checked, by matching what the scientists said about their work to their published papers.

The older kind of matching, on tables of facts with no text in them, kept improving too. Three fields (ZIP code, date of birth and sex) were enough to single out 87 percent of Americans in 1990 census data. With fifteen ordinary attributes, a 2019 study in Nature Communications found 99.98 percent of Americans could be picked out of any dataset, however much it had been trimmed.

The data this lands on

For fifteen years the loose terms have covered direct messages, forum posts and comments, dating and matrimony profiles, chat logs with AI assistants, journal entries in a mood app, and support tickets. All of it is text, written the way you write, and a dating profile in particular is long, self-describing text in your own voice by design, and that is the text the newer methods work on.

The nearest dating case is older and cruder. In 2016 two researchers published nearly 70,000 OkCupid profiles, with usernames, religion and answers to intimate questions, and defended it on the grounds that the data "was already public". Being scattered across the web gave some cover while linking it up cost money. A copy that went to a partner in 2019 as "de-identified" is still sitting wherever it went, with the same loose term on the label, and the price of putting a name back on it keeps falling.

The regulators have noticed

The UK's data protection regulator, the ICO, now tells organisations that "the more feasible and cost-effective a method becomes, the more you should consider it as a means that is reasonably likely to be used", and lists AI tools among the things to weigh. In September 2025 the EU's top court ruled that whether pseudonymised data counts as personal data depends on what the holder can reasonably do with it in the circumstances. The court applied the effort test the GDPR already had. What the tools have changed is what now counts as a reasonable effort.

What you can do

A copy that has gone out under one of those words does not come back, and nobody outside the company can check what the word covered. Three moves are left to you.

  • Ask the company what it shares and under which description. The answer is their claim, dated, in writing.
  • Object to sharing and sale where the law gives you the right, and put the objection on the record with the date.
  • Ask for deletion, and when the answer is "anonymised" or "de-identified", ask what was removed and what was kept. An anonymised record is not a deleted one, and the bar it has to clear rises every year.

Every dormant account and every fresh copy adds more text in your voice to what can be matched, and closing old accounts removes some of it.

The record of those requests is what we keep. You add the company, we work out what it likely holds and which of the loose terms sit in its policy, and we write the request, worded and ready to send. You send it from your own inbox, their reply lands in yours, and we keep the list and the dates.

Start your record →

More from the blog

ISSUE No. 03

Post 23 Sept 2026

What dating apps keep, and how the way you write can name you

A dating or matrimony profile holds your face, your faith, who you're drawn to and years of private messages. Some of that has already been shared, sold, scraped, handed to an AI company and leaked. Taking your name off doesn't hide you, because AI can now work out who wrote a text from the way it's written. Here's what these apps hold, where it has gone, and what to ask for.

ISSUE No. 01

Post 16 Jun 2026

You proved you were real. Where did the proof go?

To open an account, watch a video, or start a job, you hand your face and your ID to a company you never chose. Here is what they keep, why 'we delete it' is a claim you can't check, and what a regulator found when it looked.