What is it about?

Machine learning is increasingly used in credit scoring to predict whether borrowers are likely to repay their loans. However, credit datasets often contain far fewer “bad” borrowers than “good” borrowers, creating an imbalanced classification problem. Another challenge is reject inference: lenders usually observe repayment outcomes only for applicants who were previously approved, which can introduce sample selection bias. This review examines how the research literature has addressed these problems. We screened 490 studies from Scopus and Web of Science and analysed 88 relevant papers in detail. The findings show that many machine-learning algorithms and correction techniques have been tested across different credit datasets. SMOTE is the most frequently used approach for handling class imbalance, while robust machine-learning methods and reject-inference techniques have also been explored. The review identifies further opportunities to combine these approaches and improve the reliability of credit-scoring models.

Featured Image

Why is it important?

Credit-scoring models influence important lending decisions, but their reliability depends heavily on the data used to train them. When default cases are rare, machine-learning models may become biased toward the majority class and fail to identify risky borrowers accurately. At the same time, models trained only on previously accepted applicants can learn from a selective sample that may not represent the wider applicant population. Understanding how researchers deal with class imbalance and reject inference is therefore essential for developing more robust credit-risk models. This review brings together the main approaches used in the literature and highlights where further methodological development is still needed.

Perspectives

Better predictive algorithms alone do not solve the main difficulties of machine-learning credit scoring. The quality and structure of the training data remain fundamental. Two problems are particularly important. First, defaults are normally much less frequent than non-defaults, so conventional models may learn to favour the majority class. Second, historical lending data mainly contain repayment outcomes for applicants who were approved, making it difficult to know how rejected applicants would have behaved. Our review shows that researchers have developed many solutions, but there is still substantial room to investigate combinations of techniques, especially for imbalanced datasets. Future research should examine not only predictive accuracy but also the robustness of these methods across different datasets and lending environments.

Prof. Afshin Ashofteh
Universidade Nova de Lisboa

Read the Original

This page is a summary of: The Main Challenges of Machine Learning for Credit Scoring: A Review, January 2026, Springer Science + Business Media,
DOI: 10.1007/978-3-032-12879-9_45.
You can read the full text:

Read

Resources

Contributors

The following have contributed to this page