Generative AI & AI Assistants

The industry files

By default, your conversations can train the models, and deletion is conditional by design: safety flags override it and copies travel past the company you typed to. What was absorbed into a trained model does not come back out.

The read at a glance

Tracking priority High

They keep your conversations, which are often more candid than your search history.

If it leaks High

A leak exposes what you asked, uploaded, and confided.

Expect it kept Indefinitely

Prompts and conversations are commonly kept to "improve" the model, with no clear end date.

Identity demanded Liveness or ID

Sign-up asks for little; some add age or identity checks.

Industry profile reviewed 23 August 2026. Also machine-readable via the free API.

If it leaks

People type medical worries, legal trouble and work secrets into these tools in the first person. A transcript in your own words is far harder to explain away than a stolen password.

What repeats in the policies

The default

Your conversations train the models

The product runs on what you type. On personal plans your chats train the models unless you find the switch, and the company judges opt-outs itself. Humans read a sample of chats at several services, and a reviewed copy can stay for years after you delete.

When you delete

Deletion is conditional by design

A safety or legal flag you never see means deleted data is kept anyway, with no notice and often no end date. And once data is relabelled de-identified, no promise in the policy applies to it any more.

Who else sees it

What you typed travels past the company

Every policy read routes conversations and account data onward to contractors: hosting, moderation, support, and in places marketing. The list is given by category and left open-ended, the countries are usually unstated, and you hold no agreement with any of them.

What a company here typically holds

Worked out from the industry, not from any one company. What you actually handed over is yours to record.

Contact InfoAccount ProfileBrowsing & ActivityMessagesLocationPhotos & Biometrics Purchases · maybeHealth · maybe

What this can reveal about you

Built only from what this kind of service actually collects. A dimension that the data does not support is not listed.

Health Likely

People ask AI about symptoms and mental health frankly.

Mental health Likely

Conversations often reveal emotional state.

Political views Possible

What you ask implies your views.

What lawfully stays after you leave

Two kinds of hold. Law sets it: a statute makes them keep it. They set it: a ground the company grants itself.

Anonymised, aggregated, or AI-trained data They set it often kept indefinitely

They treat it as no longer being about you, though such data can sometimes be re-identified.

Identity and age-check records They set it as long as they choose

To prove they checked your age or identity and to block known fraud, sometimes held by a separate verification company.

Safety and abuse records They set it as long as the ban holds

To enforce bans and stop blocked or abusive users coming back.

Tax and accounting records Law sets it about 6 years

Tax and company law makes them keep billing and payment records.

Records tied to a live or potential dispute They set it the limitation period of the claim

They can keep records to defend a live or possible legal claim.

Who wants this data

Your chats are the value, and they feed the company's own models. Training is on by default, opted-in chats are kept for years, and at the big platforms assistant conversations now shape the ads you see. There are no documented sales of chat logs to brokers so far. The demand for them is in-house.

Sold or shared Likely

Conversations are valuable behavioural and intent data.

AI training High

Your prompts and uploads train the model itself unless you opt out, and opt-out works only going forward.

Even anonymised, this can still be you

Language models are the exact tool now shown to infer who wrote a text, at scale (Staab et al., ICLR 2024).

Name, date of birth, postcode Typical

Fifteen demographic attributes re-identify 99.98% of Americans in a released dataset (Rocher, Hendrickx and de Montjoye, Nature Communications, 2019); date of birth, postcode, and sex alone did it for most people in the first study of the problem (Sweeney, 2000).

Payment patterns Sometimes

Four card transactions identify 90% of people in payment data (de Montjoye et al., Science, 2015).

How you write Typical

Language models infer where a person lives, their income, and their sex from their writing alone, at near-human accuracy and at scale (Staab et al., ICLR 2024).

Face and voice Typical

A face, voice, or fingerprint template identifies a person directly; there is nothing left to anonymise, and it cannot be reissued like a password.

Browsing fingerprint Typical

Browser and device fingerprints were unique for 84% of visitors in the first large study (Eckersley, 2010), and sparse histories of what people viewed re-identified them against public reviews (Narayanan and Shmatikov, 2008).

Who you know Typical

The shape of who a person connects with re-identifies accounts across networks with no other data (Narayanan and Shmatikov, 2009).

The studies Estimating the success of re-identifications in incomplete datasets using generative models (Nature Communications 10, 3069, 2019)·Simple Demographics Often Identify People Uniquely (Carnegie Mellon University, Data Privacy Working Paper 3, 2000)·Unique in the shopping mall: On the reidentifiability of credit card metadata (Science 347 (6221), 2015)·Beyond Memorization: Violating Privacy via Inference with Large Language Models (ICLR 2024, 2024)·How Unique Is Your Web Browser? (Privacy Enhancing Technologies Symposium (PETS 2010), 2010)·Robust De-anonymization of Large Sparse Datasets (IEEE Symposium on Security and Privacy, 2008)·De-anonymizing Social Networks (IEEE Symposium on Security and Privacy, 2009)

The wording that does the work

Clauses that recur across this industry, and what each one actually permits.

“to provide and improve our services”

“to provide and improve our services”

The catch-all purpose. Analytics, profiling, personalisation and AI training all fit under it. When they want to do something new with your data, this sentence usually already allows it.

The move An objection tells them to use your data to run the service and nothing more.

“we do not sell your personal information”

“we do not sell your personal information”

Usually this means no cash changes hands. Your data can still go to ad networks, analytics firms and partners, because they count that as sharing rather than selling.

The move Use the do-not-sell switch where there is one, and put an objection in writing as well.

“service providers, partners, and affiliates”

“service providers, partners, and affiliates”

This is how your data leaves with no name attached. Recipients are described by what they do rather than named, and you cannot send a request to a company you cannot name.

The move An access request can ask for recipients by name rather than by category, and UK and EU law put that choice with you.

“aggregated or de-identified information”

“aggregated or de-identified information”

Taking your name off does not take away the pattern, and the pattern often still points at you. Policies give themselves free use of this data with no end date, on the basis that it is no longer about you.

The move If a deletion comes back as 'anonymised', keep the reply. It usually means de-identified, and it is their claim, not a fact you can check.

“retained as long as necessary, or as required by law”

“retained as long as necessary, or as required by law”

They can keep it for legal duties, tax rules, fraud prevention, possible lawsuits and their own business reasons. None of those has a firm end date, so deletion turns into something you have to argue for.

The move Which reasons apply to you, and how long each runs, is a request of its own.

“you grant us a licence to use your content”

“you grant us a licence to use your content”

This is a contract term rather than a data setting, so a privacy request cannot undo it. A careful version ends when your account does. A broad one can be passed on, never expires and survives deletion.

The move Their terms say whether the licence ends when the account does. Close the account and log the date here.

“we may use content you provide to improve our services, for example to train the models”

“we may use content you provide to improve our services, for example to train the models”

Your chats train the models by default, and improve covers research and new products. Opting out is on you, and it stops only future use, not what already went in.

The move Flip the training switch first, then put the objection in writing.

“we will delete the data within 30 days unless it is necessary to retain it for legal, compliance, or safety reasons”

“we will delete the data within 30 days unless it is necessary to retain it for legal, compliance, or safety reasons”

Every deletion has an exception. A flag you cannot see means your data stays, with no notice and no stated end date.

The move The deletion request can carry a question with it: is anything being kept.

“even if you opt out, we will use your conversations for model improvement when they are flagged for safety review”

“even if you opt out, we will use your conversations for model improvement when they are flagged for safety review”

A classifier you cannot see or contest overrides the setting you chose. A flag turns an opted-out conversation back into training material, and it is the same route that puts a human reader on it.

The move A request covers whether your conversations were flagged and what followed.

Their own policy is the one that binds them. Pin it down with a request, and keep the reply.