
Combines process mining and machine learning to detect fraud in car insurance claims. Event logs following the standard claims workflow (First Notification of Loss → Assign Claim → Claim Decision → Set Reserve → Payment Sent → Close Claim) are analyzed in Celonis to uncover bottlenecks, reworks, and suspicious deviations across thousands of records. A 21-feature dataset (categorical: case/client/policy/accident attributes; numerical: claim amounts, client age, processing duration, activity counts) feeds a Python/scikit-learn model, with SMOTE for class imbalance, to classify claims as valid, fraudulent, or violations.
Churn-risk ranking engine dealer's aftermarket business, isolating a credible 230-account cohort from a 3.1M-row export corrupted by a row-cap bug that silently dropped 63% of customers, and beating recency/frequency baselines with a bootstrap-validated 0.51 PR-AUC.
K-Means segmentation of 2,635 real customers into 4 segments (0.97 bootstrap ARI), surfacing a $668K reactivation list from a deliberately capped export.
End-to-end data mining pipeline predicting an individual's income bracket from demographic and socioeconomic data covering EDA, missing-value and outlier handling, SMOTE-balanced classes, PCA-based feature selection, and a tuned ANN benchmarked against classical ML models.