What Happens to a Document You Paste into an AI Tool
What OpenAI's and Microsoft's own enterprise privacy documentation say about training, storage, and access when you summarize a work document with AI.
Pasting a long report into an AI tool to get a quick summary feels routine now, but the document you paste in is still your organization's data, and what happens to it afterward depends entirely on which tool and which account tier you are using. The two most common workplace AI providers, OpenAI and Microsoft, both publish specifics on this, and the details matter more than a general "it's safe" or "it's risky" answer.
Training: the default differs by product tier
OpenAI's enterprise privacy page states that it does not train its models on business data by default, specifically for ChatGPT Business, ChatGPT Enterprise, ChatGPT for Healthcare, ChatGPT Edu, ChatGPT for Teachers, and the API platform, covering both inputs and outputs. This is a default, not a universal rule across every OpenAI product, which is why the specific product tier your workplace uses matters.
Microsoft's documentation for Microsoft Copilot makes a similarly specific claim: prompts, responses, and data accessed through Microsoft Graph are not used to train the foundation large language models that power Copilot. Microsoft separately notes that abuse monitoring, which includes human review of content, is available in Azure OpenAI more broadly, but Microsoft Copilot services have opted out of that human review.
Who can actually see the document
This is the detail that matters most for a document you summarize at work: can other people in your organization, or outside it, end up seeing it through the AI tool.
Microsoft's documentation states that Copilot only surfaces organizational data that a given user already has at least view permission for, using the same permission model already in place in services like SharePoint. In practice, this means Copilot summarizing a document does not bypass existing file-sharing permissions. It also means if your organization's permissions are set too broadly already (for example, a file shared with "anyone in the company" that should have been restricted), Copilot inherits that same over-sharing, since it respects whatever permission already exists rather than adding a new check.
OpenAI's documentation frames this control differently, around organizational ownership rather than permission inheritance: customers own and control their inputs and outputs where allowed by law, and control which internal data sources are connected to the account at all.
Retention: how long the document stays somewhere
OpenAI's documentation states that qualifying organizations can control how long their data is retained, specifically naming ChatGPT Enterprise, ChatGPT for Healthcare, and ChatGPT Edu as tiers where this control is available.
Microsoft's documentation describes a related but separate mechanism for Copilot: when you interact with Copilot in apps like Word, PowerPoint, Excel, or OneNote, Microsoft stores the prompt and the response as "content of interactions," encrypted while stored, and this history is visible to the user in their Copilot activity history. Admins can use Microsoft Purview to set retention policies for this data, and individual users can delete their own Copilot activity history through the Microsoft account portal.
A practical checklist before summarizing a sensitive document
Based on what both companies document about their own products, a few concrete checks are worth doing before pasting a sensitive document into any AI summarizer at work:
- Confirm which product tier your organization is actually licensed for. "We use ChatGPT" and "we use ChatGPT Enterprise" come with different default training and retention behavior according to OpenAI's own documentation.
- If you're using Microsoft Copilot, remember it surfaces only what the current user already has permission to view, so check whether your organization's existing SharePoint or file permissions are actually set correctly, rather than assuming Copilot adds its own extra restriction.
- Check whether your organization's admin has configured data retention limits; if not, assume interactions may be retained for longer than you'd expect.
- If you need the content gone afterward, check whether your tool offers a way to delete your own interaction history, as Microsoft documents for Copilot activity history.
- Treat AI-generated summaries as a draft to verify, not a finished output: Microsoft's own documentation states that generative AI responses aren't guaranteed to be 100% factual, and that users should apply judgment before sending AI output to others.
Key takeaways
- OpenAI states it does not train models on business data by default for its enterprise and business tiers, covering both inputs and outputs.
- Microsoft states that prompts, responses, and Microsoft Graph data are not used to train the foundation models behind Copilot.
- Microsoft Copilot only shows organizational data the current user already has permission to view; it does not add a new permission layer on top of your existing file-sharing settings.
- Both OpenAI and Microsoft document retention controls, but the controls differ by product tier and require action from either an admin or the individual user.
- Neither company claims AI-generated summaries are guaranteed accurate; both frame AI output as something a person should review before relying on it or sending it onward.