SSolarc Labs
Practical Article 9 min read

Redacted Documents Can Still Leak Hidden Personal Data: What to Check Before Sharing

Published August 2026 by Solarc Labs

A practical document-disclosure review for teams handling personal data: visible black boxes are not enough if underlying text, metadata, hidden content or the wrong output format can still expose information.

A document can look redacted while the underlying information remains recoverable

The ICO warns organisations to check documents for hidden personal information before disclosure and specifically addresses redaction as a source of accidental personal-data breaches. A visible black box, white text or other cosmetic change is not the same thing as removing the underlying information from the disclosed output. The safe review question is not “does this page look redacted?” It is “does the file being handed over still contain the information somewhere a recipient can reveal, copy, search, inspect or recover?”

Inspect the disclosure file, not only the authoring view

Documents can carry information outside the obvious page view: comments, tracked changes, document properties, hidden rows or sheets, layers, embedded objects and underlying text can all matter depending on the format. The ICO guidance recommends checking for hidden personal information and choosing an appropriate disclosure format rather than assuming the working document is safe because the visible page looks correct. Build the review around the exact file that will be sent or published. Re-open that output independently and test what a recipient can actually access.

Separate detection from the decision about what must be removed

Automated entity detection can help identify candidate names, email addresses, phone numbers and other personal information, but it cannot determine every contextual disclosure decision. The ICO also expects organisations to have processes for applying redactions appropriately and, in relevant access-request workflows, suitable review and authorization. Use automated detection to narrow attention, then require a human to review both the detected candidates and any sensitive information the detector may have missed.

Prefer an output where removal is irreversible for the downstream job

When the recipient only needs clean text for analysis, research or another downstream workflow, an irreversible text export can avoid some of the hidden-layer problems associated with editing the original rich document. That does not mean every disclosure should be converted to plain text: layout, evidential context and legal requirements may require a different format. The principle is narrower: the final output should contain only what the recipient needs, and the team should verify that removed information is not still present in the delivered file.

PII Redaction Desk is a bounded text-layer review workflow, not a universal disclosure engine

PII Redaction Desk supports a text-layer TXT, PDF or DOCX workflow with malware scanning, candidate PII detection, mandatory human review and irreversible clean-text export. It does not perform OCR for image-only documents, cannot guarantee automated detection finds every sensitive item and does not certify UK GDPR compliance. For a public disclosure or formal information-rights process, use the organisation’s applicable ICO/legal workflow as the authority. The product can support a bounded text-redaction job; it does not replace the disclosure decision or the review of unsupported document types.