GenAI Without Leaks
Data governance, personal information and POPIA
This module covers how to use generative AI productively without disclosing the personal information that regulation protects. It sets out the four routes by which information escapes: prompts, outputs, memorisation and training reuse. It then covers the controls that address each route, namely the difference between consumer and enterprise tools, safe prompting and redaction discipline, and the POPIA mapping across security safeguards, the operator relationship, automated decisions and cross-border transfers. The practical result is a workflow that allows an adviser to use AI on a client report lawfully.
- Explain the mechanisms by which personal information leaks into and out of generative-AI systems.
- Choose and configure tools, and write prompts, so that client and personal data is not exposed or transferred unlawfully.
- Apply POPIA sections 19, 71 and 72 and the operator concept to a real generative-AI workflow, including cross-border processing.
- Data-leakage threat modelling
- Consumer-versus-enterprise tool evaluation
- Safe prompting and redaction discipline
- POPIA operator and transborder analysis
Lessons in this module
How it lands across the four desks
Client financial data is precisely the kind of information that a careless prompt exposes, so this desk carries a high level of risk. The module gives you a workflow for producing a client report at the speed AI assistance allows, without disclosing personal information that the task never required.
You set the policy: the tool approval criteria, the operator-contract checklist, the transborder analysis, and the response when a disclosure happens anyway. The module supplies a usable version of each of those documents.
Case data is among the most sensitive information the institution holds, including alert narratives, suspicion assessments and subject identities. You learn which of it may never leave the approved environment, and the legal reasons why.
Application files contain a large volume of personal information. You learn safe-use patterns for AI-assisted underwriting work, and the leakage routes by which a drafting shortcut can become a security compromise that the institution must assess for notification.
Key literature · 7 sources
Every module rests on a verified scholarly and institutional evidence base. The full core and further reading lists open with the module.
- Carlini, N. et al. (2021) 'Extracting Training Data from Large Language Models.' USENIX Security 2021: verbatim training-data extraction demonstrated.
- Carlini, N. et al. (2019) 'The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks.' USENIX Security 2019: memorisation of rare secrets established.
- Shokri, R. et al. (2017) 'Membership Inference Attacks Against Machine Learning Models.' IEEE S&P 2017: presence-in-training-data as leakable information.
- OWASP GenAI Security Project (2025) 'OWASP Top 10 for LLM Applications': the sensitive-information-disclosure and related risk catalogue.
- NIST (2024) 'AI RMF: Generative AI Profile' (AI 600-1): data privacy as a cross-cutting generative-AI risk, with control framing.
- Protection of Personal Information Act 4 of 2013, ss 19, 20:22, 71, 72: read in full against this module's workflow mapping.