Can you put a case into that AI tool?

What 10 assistants do with what you type · checked 6 September 2026

How to read this. Of the 10 entries below, 5 were checked against the vendor's own published page and are quoted from it. 3 could not be retrieved and rest on third-party reporting. 1 is left unresolved because sources contradict each other. 1 needs no vendor at all. Every row says which it is. Checked 6 September 2026 and not maintained on a schedule: re-check before you rely on it.
How this page was made, and the conflict in it. The rows were researched and drafted with AI assistance. That is precisely why every one of them is graded and linked: the intention is that you check them, not that you trust them. This page carries a row on Anthropic's Claude and was drafted using an Anthropic model, so that is the row to verify first. Where a quote and its linked source disagree, the source is right, and we would like to be told. Published by Philip Choo, who is answerable for what it says.

Comparisons of AI tools rank them on how good they are. For anyone holding material that belongs to someone else, that is the second question. The first is whether the material may go in at all, and it is the one almost nobody answers.

It matters more than it looks. Pasting a subject's details into a hosted model is a disclosure to a third party and, where that third party is overseas, a transfer of personal data out of Singapore, which the Personal Data Protection Act 2012 does not treat as a neutral act. It may also sit badly with a confidentiality undertaking you have already signed, and with what you will later have to say about how a finding was reached. See PDPA lawful bases for investigators for the ground under that.

So this compares the boring axis: does the vendor train on what you type, how long do they keep it, and where does it go. It is also the axis that ages slowest. Prices move monthly. A vendor's position on training and residency moves rarely, and when it moves it matters.

The table

ToolTrains on your input?EvidenceWhere it is processed
Claude (Anthropic)
Commercial: API, Team, Enterprise
No, by defaultVendor page
source
“By default, we will not use your inputs or outputs from our commercial products (e.g. Claude for Work, Anthropic API, Claude Gov, etc.) to train our models.”
Not stated on this page
ChatGPT (OpenAI), API
API and platform
No, by defaultVendor page
source
“data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)”
Not stated on this page
ChatGPT (OpenAI), consumer
Free, Plus, Pro
Yes, unless you turn it offThird partyNot verified
Gemini (Google)
Gemini apps, consumer
Yes, and humans read a sampleVendor page
source
“Please don't enter confidential information that you wouldn't want a reviewer to see or Google to use to improve our services, including machine-learning technologies.”
Not stated on this page
Microsoft Copilot
Microsoft 365, commercial
NoVendor page
source
“Prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs, including those used by Microsoft Copilot.”
EU Data Boundary for EU users. "Customers outside the EU may have their queries processed in the US, EU, or other regions."
DeepSeek
All consumer use
Yes, with an opt-out rightVendor page
source
“To provide you with our services, we directly collect, process and store your Personal Data in People's Republic of China.”
People's Republic of China, stated expressly
Perplexity
Free, Pro, Max
Yes, unless you turn it offThird partyNot verified
Mistral Le Chat
Free and paid tiers
Depends on tierThird partyReported EU by default
Grok (xAI)
All
UnresolvedUnresolvedNot verified
Local open-weight models
Ollama, LM Studio, llama.cpp
No vendor receives the inputBy designYour machine
Two traps in that table. First, the API and the chat product are usually different answers for the same brand, and the brand is not the unit of analysis. Second, an opt-out is reported in several cases not to be retroactive: turning training off stops future use, it does not withdraw what has already gone.

Row by row, including what could not be checked

Claude (Anthropic) (Commercial: API, Team, Enterprise)

The same page says it covers commercial products only. The consumer plans are a separate policy and were not verified here.

Retention: Feedback you submit is kept up to 5 years, de-linked from identity. Submitting feedback is what opts that exchange in.

ChatGPT (OpenAI), API (API and platform)

The API and the chat product are different answers. Do not read one as the other.

Retention: Abuse-monitoring logs up to 30 days. Zero Data Retention exists but is subject to prior approval and excludes some endpoints.

ChatGPT (OpenAI), consumer (Free, Plus, Pro)

Reported to default to on, with the control at Settings, Data Controls, "Improve the model for everyone". openai.com and help.openai.com both refused automated retrieval, so this row is third-party reporting and is not vendor-verified. Opting out is reported not to be retroactive.

Retention: Not verified

Gemini (Google) (Gemini apps, consumer)

Google also states plainly that "a subset of chats are reviewed by human reviewers (including Google's trained service providers)". This is the clearest warning any vendor gives, and it is on the vendor's own page.

Retention: With Keep Activity on, retained until you delete. Chats selected for human review are kept up to three years even after you delete them. With it off, 72 hours.

Microsoft Copilot (Microsoft 365, commercial)

The detail most readers will miss, and it is on Microsoft's own page: "Models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary." Which model your tenant is pointed at changes the residency answer.

Retention: Stored as Copilot activity history under your tenant's own retention, manageable through Purview.

DeepSeek (All consumer use)

This is the row to read twice. The privacy policy grants a right "to opt-out of using your Personal Data for training our models", which means the default is use. Combined with the storage location, a matter with any China nexus should treat this as a red line rather than a setting.

Retention: "as long as you have an account"

Perplexity (Free, Pro, Max)

Reported: training on by default with an "AI data retention" toggle to disable it; Enterprise Pro reported excluded from training; the Sonar API reported zero-retention. perplexity.ai refused automated retrieval, so none of this is vendor-verified.

Retention: Not verified

Mistral Le Chat (Free and paid tiers)

Reported: Le Chat Team, Le Chat Enterprise and paid API plans excluded from training by default; other plans carry an opt-out toggle. The legal pages did not resolve to fetchable policy text, so this is not vendor-verified.

Retention: Not verified

Grok (xAI) (All)

Secondary sources directly contradict each other on whether Grok conversation training defaults to on or off, and x.ai refused automated retrieval. We would rather print this than pick the more confident-sounding source. Treat as unknown, which for case material means treat as no.

Retention: Not verified

Local open-weight models (Ollama, LM Studio, llama.cpp)

The only row where the answer does not depend on a policy that can change without notice. It costs capability, and for genuinely sensitive material that is usually the right trade. It also puts the retention question back where it belongs, on your own evidence handling.

Retention: Whatever your own disk does

What we could not verify, said out loud

4 of the 10 rows are not vendor-verified. openai.com, help.openai.com, perplexity.ai and x.ai all refused automated retrieval, and Mistral's legal pages did not resolve to fetchable policy text. Those rows are marked, and the marking is the point: a comparison table that looks uniformly confident is telling you less than one that admits where it thinned out.

This is the same standard as the rest of this site. The jurisdiction registry publishes seventeen open questions on the pages themselves, including the ones that will never close by research. An unmarked gap is the failure. A marked one is just the current state of the evidence.

The part that does not depend on the vendor

Every answer above can change with a policy update you will not be told about. These do not:

  • A name you were given in confidence does not need to be in the prompt. Most analytical work survives pseudonymisation completely intact. Ask the question about "the subject" and re-attach the identity yourself, afterwards, offline.
  • Decide before the matter starts, not during. The moment you are deciding whether this particular paste is acceptable is the moment you are already reasoning towards the answer you want.
  • Write down which step the model touched. If a finding is ever tested, "the tool suggested it and I checked it" is a defensible account and "I do not recall" is not. That record has to exist at the time.
  • A model may propose; a person confirms. No finding should reach a client marked confirmed because software was confident.

AI9OS maintains the Asia-Pacific Private Investigation Registry: who may investigate in each state, cited to section, dated, with the open questions published rather than hidden.

See the registry