GenAI Without Leaks
Data governance, personal information and POPIA
How to get the productivity of generative AI without leaking the personal information that regulation protects. The module shows how information actually escapes — through prompts, outputs, memorisation and training — then teaches the controls: the consumer-versus-enterprise distinction, safe prompting and redaction discipline, and the full POPIA mapping across security safeguards, the operator relationship, automated decisions and cross-border transfers. It is the module that lets an adviser safely use AI for a client report.
- Explain the mechanisms by which personal information leaks into and out of generative-AI systems.
- Choose and configure tools, and write prompts, so that client and personal data is not exposed or transferred unlawfully.
- Apply POPIA sections 19, 71 and 72 and the operator concept to a real generative-AI workflow, including cross-border processing.
- Data-leakage threat modelling
- Consumer-versus-enterprise tool evaluation
- Safe prompting and redaction discipline
- POPIA operator and transborder analysis
Lessons in this module
How it lands across the four desks
Highest stakes for you: client financial data is exactly what a careless prompt exposes. You leave with a workflow that produces the client report with AI speed and zero unlawful disclosure — the module's title promise, kept.
You set the policy: tool approval criteria, the operator-contract checklist, the transborder analysis and the incident posture when a leak happens anyway. This module hands you each artefact.
Case data is among the most sensitive information the institution holds — alert narratives, suspicion assessments, subject identities. You learn what may never leave the approved environment and why.
Application files are dense personal information. You learn safe-use patterns for AI-assisted underwriting work and the leakage routes that turn a drafting shortcut into a notifiable breach.
Key literature · 7 sources
Every module rests on a verified scholarly and institutional evidence base. The full core and further reading lists open with the module.
- Carlini, N. et al. (2021) 'Extracting Training Data from Large Language Models.' USENIX Security 2021 — verbatim training-data extraction demonstrated.
- Carlini, N. et al. (2019) 'The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks.' USENIX Security 2019 — memorisation of rare secrets established.
- Shokri, R. et al. (2017) 'Membership Inference Attacks Against Machine Learning Models.' IEEE S&P 2017 — presence-in-training-data as leakable information.
- OWASP GenAI Security Project (2025) 'OWASP Top 10 for LLM Applications' — the sensitive-information-disclosure and related risk catalogue.
- NIST (2024) 'AI RMF: Generative AI Profile' (AI 600-1) — data privacy as a cross-cutting generative-AI risk, with control framing.
- Protection of Personal Information Act 4 of 2013, ss 19, 20–22, 71, 72 — read in full against this module's workflow mapping.