AWAI at Work, Plainly
ai tools

Privacy Questions to Ask Before Using an AI Tool at Work

A checklist of privacy questions drawn from official OpenAI, Microsoft, and Google documentation to ask before adopting an AI tool at work.

Before rolling an AI assistant into daily work, whether it's for email, spreadsheets, or document review, it helps to ask a short, specific list of questions rather than relying on a general sense that "it's probably fine." The three questions below come directly from how OpenAI, Microsoft, and Google each document their own products.

Question 1: Is my data used to train the model?

This is the question most people assume the answer to without checking. The actual answer depends heavily on which product and tier you're using.

OpenAI states for its business products (ChatGPT Business, ChatGPT Enterprise, ChatGPT for Healthcare, ChatGPT Edu, ChatGPT for Teachers, and the API Platform): "We do not train our models on your data by default." The phrase "by default" is doing real work here: it implies a setting exists, and in some products a customer could in principle opt in to something different, so the baseline position, not an absolute universal rule, is what's documented.

Microsoft states, for Microsoft Copilot specifically: "Prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs, including those used by Microsoft Copilot." This is stated without a "by default" qualifier in Microsoft's current documentation, making it a firmer claim for this specific product.

Google maintains a dedicated "Generative AI in Google Workspace Privacy Hub" specifically to answer model-training and data-usage questions for Gemini in Workspace, organized under sections including "Model training and data usage" as a named topic separate from general data access questions. The existence of a dedicated, continuously updated hub page (last updated October 2026 at the time of writing) suggests this is a frequently asked question Google expects administrators to need answered in detail, rather than a single blanket statement.

Question 2: Who inside my organization can see what the AI tool sees?

This question matters more than most people realize, because an AI assistant connected to company files inherits whatever access the person using it already has, which can be broader than they think.

Microsoft's documentation states: "Microsoft Copilot only surfaces organizational data to which individual users have at least view permissions." This means Copilot will not show a user data from files they couldn't already open directly. The practical implication: if your organization's file permissions are already too loose (for example, a shared drive where everyone has view access to everything), an AI assistant does not create a new privacy problem, but it will not fix an existing one either. It surfaces exactly as much as the permission structure already allows.

Question 3: What happens to my data if I stop using the product?

Retention is the part most people never think to ask about until they're already using a tool. OpenAI's documentation specifies that customers on ChatGPT Enterprise, ChatGPT for Healthcare, and ChatGPT Edu "control how long your data is retained," which implies there is a retention period to control in the first place, distinct from the no-training-by-default commitment. A no-training promise and a no-retention promise are two separate things, and it's worth confirming both rather than assuming one implies the other.

A short checklist to bring to any AI tool evaluation

  1. Does the vendor document a "no training by default" or equivalent commitment, and does it apply to the specific product tier you're using (not just the business tier they advertise)?
  2. Can an admin control who in the organization has access, and does the tool respect existing file permissions, or does it introduce broader visibility than individual users already had?
  3. Is there a stated data retention period, and can you control it, separate from the training question?
  4. Is the commitment described with qualifiers like "by default" or stated as an absolute, and does that distinction matter for your use case (for example, handling regulated health or financial information)?
  5. Does the vendor maintain a dedicated, updated privacy resource (like Google's Privacy Hub) that you can check back on, since AI product terms change more often than most software terms of service?

Why this matters even for small teams

Smaller teams without a dedicated IT or security function are often the ones most likely to adopt an AI tool quickly, precisely because there's no formal review process slowing it down. The upside of asking these three questions yourself, even informally, is that all three answers are published directly by the vendors in the documentation already cited here. None of this requires a formal security audit. It requires reading the specific page for the specific product and tier you're actually using, rather than assuming a general reputation covers the details.

Key takeaways

  • "No training on your data" and "no data retention" are two different, separately documented commitments; check both.
  • Vendor commitments are often scoped to specific product tiers, not blanket promises across every account type the company offers.
  • AI tools connected to company files generally only see what the individual user could already access directly, so they inherit existing permission problems rather than creating new ones.
  • Watch for qualifiers like "by default" in a vendor's privacy language; they usually indicate a configurable setting exists, not an absolute rule.
  • Vendors maintain dedicated, regularly updated privacy pages for exactly these questions; read the current version for your specific product before adopting it.

Sources

  1. OpenAI, Enterprise privacy at OpenAI
  2. Microsoft Learn, Data, Privacy, and Security for Microsoft Copilot
  3. Google Workspace Help, Generative AI in Google Workspace Privacy Hub
ai toolsworkplacedata privacy