Loan analysis
Pythonscikit-learnXGBoostPandas
The problem
Which customer attributes relate to personal-loan uptake, and how do classification approaches compare?
How I solved it
Explored customer attributes, engineered and transformed features with log1p and RobustScaler preprocessing, then compared Logistic Regression, Random Forest and XGBoost with confusion matrices and class-level reports.
- EDA, feature selection and log-transformed preprocessing
- Logistic Regression, Random Forest and XGBoost comparison
- Reusable preprocessing/model artifacts and written report
The results
- Random Forest model accuracy
- 97.40%
- #1 feature in Random Forest and XGBoost
- Income
- LogReg after log transform + scaling
- +1.49 pp
Random Forest recorded 97.40% accuracy. Tree-based feature importance ranked Income first in both Random Forest (56.34%) and XGBoost (39.72%). For Logistic Regression, applying a log1p transform with RobustScaler increased recorded accuracy from 88.72% to 90.21%, a +1.49 percentage-point improvement.