A data-boundary checklist for deciding what must stay out of an AI tool, what requires formal assessment and what can be safely tested.
IF useful → CHECK boundary → HUMAN decision
A prompt can be personal data processing
Checkpoint detail
Names, email addresses, account histories, call transcripts, support tickets and identifiers do not stop being personal data when pasted into an AI service. The output may also contain personal data or infer new information. Before use, identify the purpose and lawful basis, tell people what is happening where required, and decide whether the processing is necessary. ‘It saves time’ does not remove UK GDPR duties. The organisation using the tool remains responsible for its own compliance choices even when a supplier provides the model.
Minimise before you submit
Checkpoint detail
Ask whether the task can be completed without customer data. Replace real records with synthetic examples for exploration. For an approved production use, remove fields the task does not need, reduce free text, use reference numbers instead of names where feasible, and limit the date range. Pseudonymised data can still be personal data if the organisation can reconnect it. Anonymisation requires a much stronger assessment than simply deleting a name. Data minimisation also applies to outputs, logs and evaluation sets, not only the initial prompt.
Human review does not cure every boundary problem
Checkpoint detail
Review can catch an inaccurate draft, but it cannot undo an unlawful upload, excessive retention or an unauthorised disclosure. For higher-risk uses, consider whether a data protection impact assessment is required before processing begins. Give staff an approved-tool list, examples of prohibited inputs and an incident route for accidental submission. Provide a way to correct outputs that enter customer records. Revisit the assessment when the model, provider terms, data sources or purpose changes; approval for one workflow is not blanket permission for every use.
Set a red boundary for secrets and high-risk records
Checkpoint detail
Keep passwords, access tokens, full payment-card details and security recovery information out of general AI prompts. Treat health information, biometric data, political opinions, religious beliefs, trade-union membership, sexual life or orientation, and genetic data as special category data requiring additional conditions and safeguards. Criminal-offence data has separate rules. Confidential contracts, unreleased prices and privileged legal material also need contractual and professional review even where they do not identify an individual. A consumer account is not an approved business data environment by default.
Read the supplier terms as part of the design
Checkpoint detail
Establish whether the provider acts as a processor for the use, what its contract permits, where data is stored or accessed, which subprocessors are involved, and whether prompts or outputs train models. Check retention, deletion, account controls, audit information and breach support. International transfers may require safeguards. A settings toggle is not a substitute for a processor agreement or a documented risk assessment. If the supplier cannot answer material questions, keep customer data out rather than assuming silence means safety.
Practical checklist
Customer-data boundary checklist
If any answer is unknown, pause live-data use and resolve it with the responsible data protection or legal owner.
Purpose namedCan we state the exact task and why customer data is necessary, rather than ‘using AI to improve efficiency’?
Lawful basis recordedHave we identified and documented a lawful basis, plus any Article 9 condition for special category data?
Prohibited inputs blockedDo staff know not to submit credentials, payment-card data, unapproved special category data, criminal-offence data or protected secrets?
Minimum fields onlyHave names, contact details, identifiers, dates and free text been removed or reduced wherever the task does not need them?
Test data protectedCan synthetic or properly prepared test cases replace live customer records during exploration and evaluation?
Supplier role and contract checkedDo we understand controller/processor roles, processing instructions, confidentiality, subprocessors and assistance obligations?
Training use knownDo terms and account settings clearly state whether prompts, files and outputs are used to train or improve provider models?
Retention and deletion knownCan we state how long prompts, outputs, logs and backups persist and how deletion requests are handled?
Location and transfers assessedDo we know where data is stored and accessed, and have required international-transfer safeguards been considered?
DPIA question answeredHave we screened for likely high risk and completed a DPIA before processing where required?
People informedAre privacy information, rights handling and explanations accurate for this use, including any significant automated decision-making?
Incident and review route setCan staff report accidental uploads promptly, and is reassessment triggered by changes to purpose, data, model or supplier terms?