A complete data mining pipeline built to predict income level from an individual's demographic and personal attributes. The project walks the full data science lifecycle: exploratory data analysis and visualization, missing-value and outlier handling, categorical encoding (label + ordinal), feature normalization, class-imbalance correction with SMOTE, and dimensionality reduction through PCA and feature selection. Several classical classification algorithms were trained and compared against a custom artificial neural network built with Keras/TensorFlow, with hyperparameters fine-tuned to identify the best-performing model. Results were interpreted and visualized to surface which demographic factors matter most for income prediction.
Churn-risk ranking engine dealer's aftermarket business, isolating a credible 230-account cohort from a 3.1M-row export corrupted by a row-cap bug that silently dropped 63% of customers, and beating recency/frequency baselines with a bootstrap-validated 0.51 PR-AUC.
K-Means segmentation of 2,635 real customers into 4 segments (0.97 bootstrap ARI), surfacing a $668K reactivation list from a deliberately capped export.
Combines process mining and machine learning to detect fraudulent insurance claims from event-log data.