I use AI tools a lot. And like many people, I have gradually become accustomed to typing quite a lot of information into them: ideas, drafts, questions, pieces of text and sometimes work-related information. That makes a recent piece of research highlighted by Cybernews rather uncomfortable reading.
Here is the full article: https://cybernews.com/ai-news/ai-companies-lack-disclosures-on-data-training-and-retention/. The research comes from the Production AI Institute, which created an AI Data Use Index covering 63 AI products. It looked at what companies publicly disclose about the use of customer and user data: whether data can be reused for AI training, whether users can opt out, how long information is retained and whether humans may review it. Assuming the research has been carried out properly, I find the results pretty shocking.
So what exactly happens to our data?
The important point is not that all these AI companies are secretly feeding everything we type into their models. That is not what the research says. The problem is that in many cases it is surprisingly difficult for an ordinary user to determine exactly what happens to their data. According to the researchers, only a relatively small number of products clearly state that they use customer data for training, while others offer user-controlled settings or provide information that is spread across several privacy policies, help pages and product-specific terms.
That alone is worrying. With AI, we are no longer talking about entering a few search terms into a search engine. We increasingly paste complete documents, source code, meeting notes, business plans, customer information and personal questions into these systems. The more useful AI tools become, the more information we tend to give them.
One thing that becomes clear from the research is that there is no obvious industry standard for how AI companies deal with training data. Consumer ChatGPT, for example, can use conversations to improve models depending on the user’s Data Controls settings, while OpenAI says business workspace data is not used for model training by default. Perplexity has its own settings around AI data retention, while Cursor distinguishes between users who enable Privacy Mode and those who do not. Other services, including GitHub Copilot, Slack and Notion, again apply different policies.
Your data may still be stored
So the apparently simple question, does my AI tool train on my data?, often has no simple answer. It can depend on the product, the type of account you have, the subscription you pay for, your individual privacy settings and sometimes even the specific feature you are using.
That raises another question which I find even more interesting. In cases where an AI service has a setting that allows users to opt out of data being used for model training, what exactly does that opt-out mean? Does it apply immediately? Does it also cover previously stored conversations? What happens when you provide feedback on an answer? And does opting out of training also mean your information is no longer stored, reviewed for abuse prevention or shared with third-party service providers?
These are not the same things. The Production AI Institute’s research is primarily about transparency: it examines what companies publicly say about their data practices. It is not a technical audit that verifies what actually happens inside every company’s infrastructure. That means the research does not show that companies continue using data for AI training after users have opted out. But it also does not independently prove that every opt-out is implemented exactly as users might assume.
There is another distinction that deserves much more attention: not using your data for training is not the same as not storing your data. An AI provider may promise that prompts and responses are excluded from model training while still retaining them temporarily for security, abuse monitoring, troubleshooting or other operational purposes. Enterprise and API services often have separate rules again, with different retention periods and contractual protections.
For business users in particular, asking only whether an AI company trains on your data is therefore not enough. You also need to know what information is stored, for how long, where it is processed, who can access it, which third parties may receive it and whether deleting a conversation actually removes the underlying data. You should also know whether the privacy rules for a consumer account are different from those for a business or enterprise subscription.
For years, most of us have become used to clicking through privacy notices and terms of service without paying much attention. With AI, that habit is becoming considerably more risky. Both business users and private users probably need to become much more conscious about the services they use and the information they provide.
We need to pay more attention
Before using an AI tool for anything sensitive, it makes sense to check its privacy and data controls. Is model improvement enabled by default? Is there an opt-out? Does the opt-out apply to all data or only certain types of conversations? Are consumer and business accounts treated differently? And perhaps most importantly, what exactly does the company mean when it says it does not use your data for training?
I would like AI companies to make this much simpler. Every AI service should have one clearly visible page explaining in plain language what happens to everything you submit: whether it is used for training, whether that is opt-in or opt-out, how long data is retained, whether humans can review it, whether third parties receive it and whether you can permanently delete it.
That information should not require digging through five different policy documents. Because if this research is even broadly representative of the market, we have probably been paying far too little attention to a very basic question: What actually happens to our data after we press Enter?