Predicting Customer Purchase Behavior


End-to-End ML Project

This is a complete machine learning case study using real-world behavioral data. It covers ETL, EDA, Feature Engineering, Model Training, Hyperparameter Tuning(GridSearchCV), and performance evaluation..


Data Source:

The dataset used in this project was publicly available on Kaggle:
Predict Customer purchase Behavior Dataset : Kaggle Dataset by Rabie El Kharoual

Dataset Description: This dataset contains information on customer purchase behavior across various attributes, aiming to help data scientists and analysts understand the factors influencing purchase decisions. The dataset includes demographic information, purchasing habits, and other relevant features.

Features:

  • Age: Customer's age
  • Gender: Customer's gender (0: Male, 1: Female)
  • Income: Annual income of the customer in dollars
  • Number of Purchases: Total number of purchases made by the customer
  • Product Category: Category of the purchased product (0: Electronics, 1: Clothing, 2: Home Goods, 3: Beauty, 4: Sports)
  • Time Spent on Website: Time spent by the customer on the website in minutes
  • Loyalty Program: Whether the customer is a member of the loyalty program (0: No, 1: Yes)
  • Discounts Availed: Number of discounts availed by the customer (range: 0-5)
  • PurchaseStatus (Target Variable): Likelihood of the customer making a purchase (0: No, 1: Yes)

Target Variable: Distribution of the Target Variable (PurchaseStatus):

  • 0 (No Purchase): 48%
  • 1 (Purchase): 52%

Phase 1: Data Cleaning

I began by importing the raw dataset and understanding its structure. After confirming that there were no missing values, I identified and converted several numerical columns to categorical type (e.g. `gender`, `LoyaltyProgram`, `ProductCategory` and `PurchaseStatus`). Finally, the cleaned dataset was saved for reuse in the future steps.
Link to ipynb Notebook with visualization: 01_data_cleaning.ipynb

Phase 2: Exploratory Data Analysis (EDA)

In this phase, I performed univariate and bivariate analysis to explore patterns and relationships in the data.

Univariate Analysis

  • TimeSpentOnWebsite: Slightly positive trend; peak en engagement near 45 minutes.
  • AnnualIncome : Surprisingly balanced; no transformation needed.
  • NumberOfPurchases : Spike at 20 suggests a possible system cap.
  • DiscountsAvailed : Fairly even distribution across all bins.
  • Age : Bimodal pattern with peaks in younger and older ranges.

Bivariate Analysis(vs. `PurchaseStatus`)

  • TimeSpentOnWebsite : Clear increase in median time for purchasers.
  • AnnualIncome : Purchasers had notably higher income.
  • DiscountsAvailed : Slightly more discounts used by purchasers.

Link to ipynb Notebook: 02_exploratory_data_analysis.ipynb

Phase 3: Modeling and Evaluation

After preparing the features and encoding categorical variables, I split the dataset and trained a baseline RandomForestClassifier

Model F1-Score Accuracy Recall Precision
Random Forest (Baseline Model) 0.90 0.92 0.88 0.93
Random Forest (Tuned Model) 0.91 0.93 0.88 0.94

The baseline model acheived(purchase class):

  • F1-Score: 0.90.
  • Accuracy : 92%.

I then used GridSearchCV to optimize the key hyperparameters such as n_estimators, max_depth, min_samples_split, and min_samples_leaf.
The tuned model imporved performance further:

  • F1-Score: 0.91.
  • Accuracy : 93%.
  • Precision : 0.94

This confirmed that the featrures were meaningful and the model generalized well.
I would consider this version production-ready for business insights or deployment

Link to ipynb Notebook: 03_modeling_pipeline.ipynb