Can I Upload Client Documents to an AI Tool? Minimise the Personal Data Before You Send Anything
Published August 2026 by Solarc Labs
Start with “what does the task need?” — not “can the model read the whole file?”
A client document can contain names, addresses, account details, signatures, staff information, third-party data and context that has nothing to do with the AI task. Before uploading or pasting anything, define the exact purpose and the minimum information needed to achieve it. The ICO describes data minimisation as keeping personal data adequate, relevant and limited to what is necessary. Its current AI audit guidance also tells organisations to review relevance across AI processing and specifically lists document cropping or redaction as an option to consider for collection and sharing.
Redaction is only one control; it does not answer permission, contract or confidentiality questions
Removing unnecessary identifiers can reduce the amount of personal information exposed to a downstream system, but it does not automatically make the remaining use acceptable. Check the organisation’s purpose, lawful basis where required, confidentiality obligations, customer commitments, vendor terms, security review, retention expectations and any sector-specific restrictions that apply to the real workflow. If the organisation has not approved the AI service for the intended data, a redacted file should not be treated as a workaround for that governance decision.
Decide which details the downstream task can work without
For a drafting or classification task, a person’s full name, complete address, account number or exact date may be irrelevant even though the surrounding business text is useful. Replace or remove information only when the task can still be performed correctly without it. Do not mechanically strip every identifier if doing so destroys meaning the authorised task genuinely requires. Data minimisation means using the minimum appropriate information for the purpose, not blindly deleting context until the document becomes misleading.
Automated PII detection is a candidate finder, not proof that the document is clean
Names and structured identifiers can be detected automatically, but sensitive information also appears in unusual formats, free text and context that a detector can miss or misclassify. A high-confidence automated result should therefore remain a proposal for review rather than a certification that no personal data remains. For consequential client material, inspect the proposed redactions, add missed spans, reject false positives and review the final clean output before it is released to another system.
Keep unsupported files out of the “probably fine” path
A scanned image-only PDF is not the same thing as a document with a usable text layer. If a redaction workflow does not support OCR, it should say so and stop rather than silently passing an unreadable page as clean. The same principle applies to embedded objects, unusual document structures and any state the tool cannot reliably extract and review. An explicit unsupported result is safer than a green-looking report based on content the system never actually inspected.
Review the clean export separately from the original
After redaction, inspect what will actually be handed to the downstream AI or analysis process. Confirm that required meaning remains, selected sensitive spans are no longer present in the clean output, and the audit summary does not itself repeat the raw personal data. Keep access to the original restricted according to the organisation’s real retention and security rules. The goal is a controlled handoff: only the information intended for the downstream task crosses that boundary.
PII Redaction Desk is a review workflow, not a GDPR or “safe for AI” certificate
PII Redaction Desk supports text-layer TXT, PDF and DOCX in its current bounded workflow. It uses automated candidate detection, requires human review and produces a clean-text export after approved redactions. The current pilot explicitly makes no OCR claim and no GDPR/compliance certification claim. Automated detection can miss sensitive information. The operator remains responsible for reviewing the document, deciding what the downstream task is authorised to process and applying the organisation’s legal, contractual, security and retention requirements.
Primary sources
Sources used for this article
Continue the job