← All projectsData Science/Analytics

Loan analysis

Pythonscikit-learnXGBoostPandas

The problem

Which customer attributes relate to personal-loan uptake, and how do classification approaches compare?

How I solved it

Explored customer attributes, engineered and transformed features with log1p and RobustScaler preprocessing, then compared Logistic Regression, Random Forest and XGBoost with confusion matrices and class-level reports.

  • EDA, feature selection and log-transformed preprocessing
  • Logistic Regression, Random Forest and XGBoost comparison
  • Reusable preprocessing/model artifacts and written report

The results

Random Forest model accuracy
97.40%
#1 feature in Random Forest and XGBoost
Income
LogReg after log transform + scaling
+1.49 pp

Random Forest recorded 97.40% accuracy. Tree-based feature importance ranked Income first in both Random Forest (56.34%) and XGBoost (39.72%). For Logistic Regression, applying a log1p transform with RobustScaler increased recorded accuracy from 88.72% to 90.21%, a +1.49 percentage-point improvement.

πŸ€—