Model Risk and Explainability
Interrogating the black box
The discipline of not trusting a model you cannot interrogate. It introduces model-risk management in the lineage of supervisory guidance, the main explainability methods and their genuine limits, and the difference between an explanation that satisfies a regulator and one that satisfies an engineer — framed for a South African institution and its Prudential Authority expectations. It closes the Core certificate by giving every prior module its governance spine: the verification routine, the monitoring stack and the fairness audit all assume a professional who can read, stress and challenge a model's documentation. This module builds that professional.
- Explain what model risk is and set out a validation and governance lifecycle for a deployed model.
- Use and critique post-hoc explanation methods, and know when an inherently interpretable model is the better choice.
- Judge whether a model's explanation is adequate for the decision it supports and the duty attached to it.
- Model-risk framing
- Reading and stress-testing model documentation
- Interpreting explainability outputs and their failure modes
- The human-oversight posture the EU AI Act requires
Lessons in this module
How it lands across the four desks
You own the framework: the model inventory, the validation cycle, effective challenge, and the documentation standard a supervisor samples. This module gives you the vocabulary and the checklist.
You apply the framework to the scoring models of Module 5 — and gain the explainability literacy to turn feature contributions into the lawful adverse-action reasons that module demanded.
You apply it to the detection models of Module 4: drift monitoring is why your typology feedback matters, and explanation outputs are what turn a score into an investigable alert.
You learn what a defensible model explanation sounds like to a client — and what questions to ask before your practice relies on any scoring or recommendation engine.
Key literature · 6 sources
Every module rests on a verified scholarly and institutional evidence base. The full core and further reading lists open with the module.
- Board of Governors of the Federal Reserve System / OCC (2011) 'SR 11-7: Supervisory Guidance on Model Risk Management' — the founding framework: the two-limb definition, the lifecycle, effective challenge.
- Ribeiro, M. T., Singh, S. & Guestrin, C. (2016) '"Why Should I Trust You?" Explaining the Predictions of Any Classifier (LIME).' KDD 2016 — local model-agnostic explanation and its mechanism.
- Lundberg, S. M. & Lee, S.-I. (2017) 'A Unified Approach to Interpreting Model Predictions (SHAP).' NeurIPS 2017 — the Shapley-value attribution framework that dominates practice.
- Rudin, C. (2019) 'Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.' Nature Machine Intelligence 1(5) — the load-bearing counter-position and this module's decision rule.
- Barredo Arrieta, A. et al. (2020) 'Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI.' Information Fusion 58 — the field survey framing the toolbox and its trade-offs.
- SARB Prudential Authority & FSCA (2025) 'Artificial Intelligence in the South African Financial Sector' — the explainability-methods survey data, the governance-gap findings, and the regulatory-constraint rankings cited in this module.