Credit Risk Under Regime Change
This project investigates how credit risk models behave when the environment they were trained on changes. Using LendingClub loan data from 2007 to 2015, the objective was not only to predict default, but to examine how model reliability changes when relationships between borrower characteristics and outcomes shift over time.
Logistic regression models were evaluated with and without regularisation, while random forest hyperparameters were tuned using Bayesian optimisation. Both approaches used borrower and loan characteristics including credit grade, interest rate, debt-to-income ratio, income, loan amount, and term. Rather than relying on a conventional random train-test split, performance was evaluated through a rolling time-based framework, with models trained on recent historical data and tested over progressively later periods.
Key components include:
- Designing a time-aware evaluation framework to assess model reliability under distribution shift
- Comparing linear and tree-based models in terms of predictive performance, stability, and calibration
- Evaluating how predicted default probabilities align with realised outcomes over time
- Tracking shifts in feature importance to identify changing drivers of risk
The models behaved differently as conditions changed. Random forests achieved stronger performance during relatively stable periods, but their AUC and calibration deteriorated more sharply as the data moved further from the conditions on which they were trained. Logistic regression achieved lower peak performance, but remained more stable over time.
Feature importance also changed across periods, particularly for variables such as loan grade and interest rate. This suggests that the relationships used by a model to assess risk are not fixed, and that strong historical performance does not necessarily imply reliable future behaviour.
The project highlights the importance of evaluating models in the context in which they will actually be used. When underlying relationships change, robustness and calibration can matter more than maximising predictive performance on historical data.