Model Risk and Explainability
Interrogating the black box
This module is about working responsibly with a model whose internal reasoning you cannot inspect directly. It introduces model-risk management as it developed from international supervisory guidance, the main explainability methods and their limits, and the difference between an explanation that satisfies a regulator and one that satisfies an engineer, all framed for a South African institution and the expectations of the Prudential Authority. It closes the Core certificate by setting out the governance framework the earlier modules assume. The verification routine, the monitoring stack and the fairness audit each depend on professionals who can read a model's documentation, test its claims and question its conclusions, and this module develops those skills.
- Explain what model risk is and set out a validation and governance lifecycle for a deployed model.
- Use and critique post-hoc explanation methods, and know when an inherently interpretable model is the better choice.
- Judge whether a model's explanation is adequate for the decision it supports and the duty attached to it.
- Model-risk framing
- Reading and stress-testing model documentation
- Interpreting explainability outputs and their failure modes
- Applying the human-oversight standard set out in the EU AI Act
Lessons in this module
How it lands across the four desks
You are responsible for the framework itself: the model inventory, the validation cycle, effective challenge, and the documentation standard a supervisor will sample. This module supplies the vocabulary and a working checklist.
You apply the framework to the scoring models of Module 5, and you learn to read explainability outputs well enough to turn feature contributions into the lawful adverse-action reasons that module required.
You apply it to the detection models of Module 4. Drift monitoring is the reason your typology feedback matters, and explanation outputs are what turn a score into an alert an investigator can act on.
You learn what a defensible model explanation sounds like to a client, and what to ask before your practice relies on any scoring or recommendation engine.
Key literature · 6 sources
Every module rests on a verified scholarly and institutional evidence base. The full core and further reading lists open with the module.
- Board of Governors of the Federal Reserve System / OCC (2011) 'SR 11-7: Supervisory Guidance on Model Risk Management': the founding framework: the two-limb definition, the lifecycle, effective challenge.
- Ribeiro, M. T., Singh, S. & Guestrin, C. (2016) '"Why Should I Trust You?" Explaining the Predictions of Any Classifier (LIME).' KDD 2016: local model-agnostic explanation and its mechanism.
- Lundberg, S. M. & Lee, S.-I. (2017) 'A Unified Approach to Interpreting Model Predictions (SHAP).' NeurIPS 2017: the Shapley-value attribution framework that dominates practice.
- Rudin, C. (2019) 'Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.' Nature Machine Intelligence 1(5): the principal counter-position and this module's decision rule.
- Barredo Arrieta, A. et al. (2020) 'Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI.' Information Fusion 58: the field survey framing the toolbox and its trade-offs.
- SARB Prudential Authority & FSCA (2025) 'Artificial Intelligence in the South African Financial Sector': the explainability-methods survey data, the governance-gap findings, and the regulatory-constraint rankings cited in this module.