The other day, I decided to Google my credit history. Just for general knowledge. I won’t provide any links here since this isn’t an advertising post.
Well, I was surprised to find that my credit score is 939 out of 1000 (this isn’t the scoring that banks calculate internally, but an external credit score from a bureau). And considering I have a history of multiple loans, credit cards, and even several late payments, who knows what I’ve done in my “credit life.” However, the score turned out to be quite decent.

And I thought it was time to dust off my credit ML project and show how credit scoring works using an example (the fully functional pipeline is here: https://github.com/John-Gear/Home-Credit-Default-Risk-Kaggle-)
What is Credit Scoring Anyway?
Imagine a few people come to you asking for money:
- you know that Masha will definitely pay you back
- Dima is questionable, but you might lend him some
- Misha definitely won’t pay you back since he has already borrowed money from Masha (p.s. No hard feelings, Misha)
When there are only a few borrowers, you can evaluate them based on intuition.
But what if a bank has tens of thousands or hundreds of thousands of borrowers?
And Here Comes Machine Learning
Here’s how it works in reality using data from Home Credit clients (p.s. all data is anonymized, we’re not violating anyone’s rights):
I took a real dataset from Home Credit (over 300,000 clients) and built a model (a machine learning algorithm) based on it that:
- looks at the client (income, history, loans, etc. — about 169 features in total)
- outputs the probability of default (for example: 0.23 = 23% chance of not paying back)
- then based on this probability, we set a threshold: for instance, if the probability of default is above 30% — we don’t give a loan, if it’s below — we do
Why default specifically? Because we’re primarily interested in whether the person will repay the loan.
What’s the Key Idea?
In short, the key idea is not about “accuracy for the sake of accuracy.” It’s about the bank’s bottom line.
In credit scoring, the logic is simple:
- if the client is good → we earn
- if bad → we lose the principal and miss out on interest
Now the task is to find a threshold that translates probability into a “give / not give” decision, maximizing the bank’s overall profit.
Simply put — it’s the criterion that translates our “Misha” into a decision: to give a loan or not.
How It Looks in Practice
When there are 300,000 clients coming in daily for loans — this needs to be automated. That’s what scoring is responsible for.
Here’s an example from my project.
The model assesses risk:
- client A → 5% risk of default → we give a loan
- client B → 40% risk of default → we don’t give
But where’s the boundary?
I found the optimal threshold through a series of tests:
- we approve ~83% of clients
- we identify ~48% of defaults (recall)
- result: ≈ 514 million rubles expected profit (a forecast I made up)
Yes, there are losses. Yes, there are false approvals.
And yes, there are people who were denied but are actually trustworthy
(for example, someone got a new high-paying job and quit drinking/smoking/moved out — just kidding).
But overall, it’s balanced: with this approach, the bank consistently earns whether it has 300,000 clients or 3 million.
This Is No Longer a Model for Metrics — It’s Business
If you’re interested, here’s what’s inside the ML part:
CatBoost as the Base Model
The very algorithm that analyzes the behavior of 300,000 clients.
It works very well out of the box, thanks to the guys at Yandex.
Comparison with LightGBM
A comparison of the effectiveness of predicting defaults.
CatBoost (Yandex) vs LightGBM (Microsoft).
K-Fold CV + OOF Validation
We split the data into parts, train on some, and test on others, and so on.
OOF (out-of-fold) means that each record has been in the test at least once, not in training.
The goal is to understand that the models didn’t just “get lucky,” but are genuinely capturing patterns.
In the end, we get a model with tuned hyperparameters.
Threshold Tuning
This is where we choose that very threshold that translates probability into the final “give / not give” decision.
The Most Important Insight
Credit scoring is not about model accuracy (as you’ve already understood).
It can make mistakes in your specific case. For example, when you went to get a loan for a new G-Class after a raise and got denied — I get it, that hurts.

But it is optimal on average across all bank clients.
Credit scoring is about the bank’s money.
If we find most clients who will actually pay back — we’re in the black.
Even if there are individual mistakes and “Misha” is actually a decent borrower.
To Simplify It All
Credit scoring is a model that decides who to lend money to, so that at the end of the year, the bank maximizes its profit.