Can you put a case into that AI tool?
Comparisons of AI tools rank them on how good they are. For anyone holding material that belongs to someone else, that is the second question. The first is whether the material may go in at all, and it is the one almost nobody answers.
It matters more than it looks. Pasting a subject's details into a hosted model is a disclosure to a third party and, where that third party is overseas, a transfer of personal data out of Singapore, which the Personal Data Protection Act 2012 does not treat as a neutral act. It may also sit badly with a confidentiality undertaking you have already signed, and with what you will later have to say about how a finding was reached. See PDPA lawful bases for investigators for the ground under that.
So this compares the boring axis: does the vendor train on what you type, how long do they keep it, and where does it go. It is also the axis that ages slowest. Prices move monthly. A vendor's position on training and residency moves rarely, and when it moves it matters.
The table
| Tool | Trains on your input? | Evidence | Where it is processed |
|---|---|---|---|
| Claude (Anthropic) Commercial: API, Team, Enterprise | No, by default | Vendor page source “By default, we will not use your inputs or outputs from our commercial products (e.g. Claude for Work, Anthropic API, Claude Gov, etc.) to train our models.” | Not stated on this page |
| ChatGPT (OpenAI), API API and platform | No, by default | Vendor page source “data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)” | Not stated on this page |
| ChatGPT (OpenAI), consumer Free, Plus, Pro | Yes, unless you turn it off | Third party | Not verified |
| Gemini (Google) Gemini apps, consumer | Yes, and humans read a sample | Vendor page source “Please don't enter confidential information that you wouldn't want a reviewer to see or Google to use to improve our services, including machine-learning technologies.” | Not stated on this page |
| Microsoft Copilot Microsoft 365, commercial | No | Vendor page source “Prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs, including those used by Microsoft Copilot.” | EU Data Boundary for EU users. "Customers outside the EU may have their queries processed in the US, EU, or other regions." |
| DeepSeek All consumer use | Yes, with an opt-out right | Vendor page source “To provide you with our services, we directly collect, process and store your Personal Data in People's Republic of China.” | People's Republic of China, stated expressly |
| Perplexity Free, Pro, Max | Yes, unless you turn it off | Third party | Not verified |
| Mistral Le Chat Free and paid tiers | Depends on tier | Third party | Reported EU by default |
| Grok (xAI) All | Unresolved | Unresolved | Not verified |
| Local open-weight models Ollama, LM Studio, llama.cpp | No vendor receives the input | By design | Your machine |
Row by row, including what could not be checked
Claude (Anthropic) (Commercial: API, Team, Enterprise)
The same page says it covers commercial products only. The consumer plans are a separate policy and were not verified here.
Retention: Feedback you submit is kept up to 5 years, de-linked from identity. Submitting feedback is what opts that exchange in.
ChatGPT (OpenAI), API (API and platform)
The API and the chat product are different answers. Do not read one as the other.
Retention: Abuse-monitoring logs up to 30 days. Zero Data Retention exists but is subject to prior approval and excludes some endpoints.
ChatGPT (OpenAI), consumer (Free, Plus, Pro)
Reported to default to on, with the control at Settings, Data Controls, "Improve the model for everyone". openai.com and help.openai.com both refused automated retrieval, so this row is third-party reporting and is not vendor-verified. Opting out is reported not to be retroactive.
Retention: Not verified
Gemini (Google) (Gemini apps, consumer)
Google also states plainly that "a subset of chats are reviewed by human reviewers (including Google's trained service providers)". This is the clearest warning any vendor gives, and it is on the vendor's own page.
Retention: With Keep Activity on, retained until you delete. Chats selected for human review are kept up to three years even after you delete them. With it off, 72 hours.
Microsoft Copilot (Microsoft 365, commercial)
The detail most readers will miss, and it is on Microsoft's own page: "Models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary." Which model your tenant is pointed at changes the residency answer.
Retention: Stored as Copilot activity history under your tenant's own retention, manageable through Purview.
DeepSeek (All consumer use)
This is the row to read twice. The privacy policy grants a right "to opt-out of using your Personal Data for training our models", which means the default is use. Combined with the storage location, a matter with any China nexus should treat this as a red line rather than a setting.
Retention: "as long as you have an account"
Perplexity (Free, Pro, Max)
Reported: training on by default with an "AI data retention" toggle to disable it; Enterprise Pro reported excluded from training; the Sonar API reported zero-retention. perplexity.ai refused automated retrieval, so none of this is vendor-verified.
Retention: Not verified
Mistral Le Chat (Free and paid tiers)
Reported: Le Chat Team, Le Chat Enterprise and paid API plans excluded from training by default; other plans carry an opt-out toggle. The legal pages did not resolve to fetchable policy text, so this is not vendor-verified.
Retention: Not verified
Grok (xAI) (All)
Secondary sources directly contradict each other on whether Grok conversation training defaults to on or off, and x.ai refused automated retrieval. We would rather print this than pick the more confident-sounding source. Treat as unknown, which for case material means treat as no.
Retention: Not verified
Local open-weight models (Ollama, LM Studio, llama.cpp)
The only row where the answer does not depend on a policy that can change without notice. It costs capability, and for genuinely sensitive material that is usually the right trade. It also puts the retention question back where it belongs, on your own evidence handling.
Retention: Whatever your own disk does
What we could not verify, said out loud
4 of the 10 rows are not vendor-verified. openai.com, help.openai.com, perplexity.ai and x.ai all refused automated retrieval, and Mistral's legal pages did not resolve to fetchable policy text. Those rows are marked, and the marking is the point: a comparison table that looks uniformly confident is telling you less than one that admits where it thinned out.
This is the same standard as the rest of this site. The jurisdiction registry publishes seventeen open questions on the pages themselves, including the ones that will never close by research. An unmarked gap is the failure. A marked one is just the current state of the evidence.
The part that does not depend on the vendor
Every answer above can change with a policy update you will not be told about. These do not:
- A name you were given in confidence does not need to be in the prompt. Most analytical work survives pseudonymisation completely intact. Ask the question about "the subject" and re-attach the identity yourself, afterwards, offline.
- Decide before the matter starts, not during. The moment you are deciding whether this particular paste is acceptable is the moment you are already reasoning towards the answer you want.
- Write down which step the model touched. If a finding is ever tested, "the tool suggested it and I checked it" is a defensible account and "I do not recall" is not. That record has to exist at the time.
- A model may propose; a person confirms. No finding should reach a client marked confirmed because software was confident.
AI9OS maintains the Asia-Pacific Private Investigation Registry: who may investigate in each state, cited to section, dated, with the open questions published rather than hidden.
See the registry