SSolarc Labs
Practical Article 10 min read

Enterprise Customer Asks “Do You Train AI on Our Data?” How Should a SaaS Vendor Answer?

Published September 2026 by Solarc Labs

Do not collapse product behavior, upstream model-provider terms, fine-tuning, evaluation, feedback, retention and vague “service improvement” into one yes/no. Map the exact customer-data path, state what each processor may do today, cite the controlling evidence and leave unsupported states explicit.

Split “do you train on our data?” into the processing jobs the buyer is actually asking about

A security questionnaire may use one row for several materially different behaviors: training a new model, fine-tuning an existing model, evaluating model quality, retaining prompts or outputs, using feedback for improvement, creating embeddings or caches, and sending data to an upstream model API. A defensible answer says which of those exist in the purchased service and which do not. NCSC secure-AI guidance explicitly distinguishes training a new model, using an existing model, fine-tuning and accessing a model through an external API. It also tells providers to consider user feedback or continuous-learning data as part of the training-data and supply-chain story. That is why one broad “we do not train on customer data” sentence can be misleading when other improvement or provider-side processing still exists.

Separate your own product behavior from the upstream model provider

If your SaaS calls an external model API, the buyer needs two answers: what your application does with customer prompts, files and outputs, and what the external model provider is permitted to do with the data it receives through the specific service path you use. NCSC recommends due diligence on external model providers and appropriate controls on data sent outside your organisation. Do not inherit a provider marketing statement from a different consumer product, plan or API route. Record the provider, service or API in scope, the applicable terms or enterprise setting, the evidence date and the owner responsible for re-checking it after a material provider or plan change.

Keep training, fine-tuning, evaluation, feedback and generic “service improvement” separate

A product may truthfully say that customer prompts are not used to train a foundation model while still retaining limited data for abuse review, evaluation, debugging, support or another documented purpose. Conversely, a feedback control may create a separate path into evaluation or model improvement. Use the actual contract, privacy information, product settings and engineering data flow to name these purposes instead of treating every purpose as either “training” or “not training.” ICO guidance treats personal data in training data, fine-tuning data, model inputs and related AI lifecycle processing as data-protection-relevant. It also requires transparency about the purposes for processing personal data, retention and who the data is shared with. Those are separate facts that should survive procurement review independently.

If personal data can become training data, the answer needs a real transparency and rights owner

ICO guidance says people must be informed if their personal data is going to be used to train an AI system, and that data-protection rights can apply to personal data in training datasets, deployment inputs and in some cases the model itself. A vendor questionnaire is not a substitute for those obligations, but it should not contradict the privacy information or internal processing record that owns them. If the service does not permit customer content to be used for training under the applicable contract, cite that evidence. If some opt-in research, feedback or fine-tuning path exists, describe the trigger and scope rather than answering “never” across every possible configuration.

Make settings, plan differences and opt-in states part of the evidence

A no-training control that depends on a workspace setting, account tier, API configuration or separate enterprise term is only as strong as the state that actually applies to the customer. Record whether the protection is contractual, configuration-based or both; who can change the setting; and whether a provider change could alter the answer. NCSC deployment guidance asks AI providers to be transparent about where and how user data might be used, accessed or stored, including whether it is used for model retraining or reviewed by employees or partners. Preserve that same precision in the buyer answer instead of promoting a configurable state into a universal product promise.

Use an evidence packet that can be refreshed when the AI supply chain changes

For each material AI path, record the customer data category, purpose, internal processing, external model or provider, retention or logging state, training/fine-tuning/evaluation state, controlling contract or policy, relevant configuration, evidence owner and last review date. Treat unknown as unknown until the responsible owner verifies it. When the model provider, model family, product plan, feedback path, retention setting or contract changes, trigger a review of the approved questionnaire answer. The goal is not a permanent sentence; it is a current answer that can be traced back to current evidence.

VendorOS can organise the approved answer; it cannot manufacture a no-training guarantee

VendorOS Security Questionnaire Rescue can map this buyer question to approved wording, current source evidence, owners, review dates, customer-safe exposure and explicit gaps. It can keep repeated questionnaires from drifting into contradictory answers. It does not inspect every upstream provider automatically, alter a provider training policy, create a lawful basis, amend a DPA, prove that deleted data never existed in logs or certify AI/data-protection compliance. If the evidence does not support a requested no-training or zero-retention commitment, preserve that as a deal constraint for the responsible product, security, privacy or legal owner.