STEG Fraud Detection in Electricity & Gas

Detecting fraudulent meter manipulation for Tunisian utility STEG. Given 15 years of client billing history, predict which clients are committing fraud at a 5.6% positive rate. Key innovations: stochastic feature selection across 100 LightGBM iterations and a 14-model diversity ensemble with an elastic net meta-learner.

🏆 Rank 6 / 191 competitors 20 continuous hours Only 54 succeeded to submit LightGBM · XGBoost · CatBoost · H2O

TL;DR: Feature interaction spaces are too large for greedy search. Stochastic sampling across many random subsets explores combinations that sequential selection structurally cannot find — and in time-constrained settings, it's the only viable path to the global optimum.

The Problem

The first thing I noticed was the imbalance: only 5.6% of records were fraudulent. STEG, the Tunisian electricity and gas utility, was losing 200 million Dinars to meter manipulation — but in the data, fraudsters were needles in a haystack. Given 15 years of billing history for each client (invoices, counter readings, consumption patterns from 2005 to 2019), I needed to identify which clients were cheating.

Three challenges made this harder than a standard binary classification. First, the data was invoice-level — multiple rows per client — so I had to engineer meaningful aggregations before any model could run. Second, categorical variables had inconsistent factor levels between train and test, meaning naive encoders would silently break on unseen values. Third, with only 20 continuous hours for the hackathon, every design decision had to be high-impact and fast to implement.

The competition attracted 191 teams, but only 54 managed to submit a valid entry within the time constraint. This told me the data wrangling alone was a significant filter — most teams got stuck on the engineering before reaching the modeling.

🏆
Rank 6 out of 191 competitors — in 20 continuous hours Only 54 teams managed to submit a valid entry in the hackathon timeframe.

My Approach

I knew that with a 20-hour constraint, I couldn't afford to spend time on approaches that might not pay off. So I focused on two high-leverage strategies: stochastic feature selection and ensemble diversity. Standard greedy feature selection (forward or backward) gets trapped in local optima — it evaluates features one at a time and misses combinations that are only useful together. Instead, I ran 100 iterations where each randomly sampled 12 to 35 features, trained a LightGBM, and recorded the AUC. The features appearing in the top-5 performing subsets formed my final feature set.

For the high-cardinality region variable, I applied five different encoding strategies and kept all of them as separate features: target mean, Weight of Evidence, James-Stein shrinkage, M-estimator, and leave-one-out. Rather than choosing one encoding before seeing the data, I let the model learn which encoding captures the most signal at each tree split. This is a small trick that adds almost no engineering time but consistently improves gradient-boosted models on high-cardinality categoricals.

The ensemble combined 14 base models across three diversity axes: algorithm (LightGBM, XGBoost, CatBoost, H2O AutoML, Random Forest), hyperparameters (deep vs. shallow trees, high vs. low learning rate, dart vs. gbdt boosters), and random seed. An elastic net meta-learner trained on the 14 probability outputs learned the optimal blend weight for each model — favoring those that were both accurate and contributed unique signal that other models missed.

For the class imbalance, I tested multiple strategies: stratified undersampling, SMOTE, 50/50 sampling for fast iteration, and full data with class weighting for final models. The answer depended on the algorithm — tree models preferred weighted full data while linear models benefited from rebalancing.

Key Decisions

Stochastic Feature Selection (100 Random Subsets)

Instead of evaluating features sequentially, I sampled random combinations and let the results reveal which features work together. This explores the interaction space that greedy search structurally cannot reach — and in a hackathon, it runs in parallel while I work on other things.

Five Encoding Strategies as Parallel Features

For high-cardinality categoricals, I computed target mean, WOE, James-Stein, M-estimator, and leave-one-out — all kept as separate columns. The model itself decides which encoding is most informative per split. No upfront commitment to a single encoding strategy.

Level Harmonization Before Encoding

Train and test had mismatched factor levels. Rather than dropping unseen levels, I merged rare categories into meaningful groups based on their fraud rate. This preserves signal while ensuring the pipeline handles unseen values gracefully.

14-Model Diverse Ensemble with Elastic Net Meta-Learner

Diversity across algorithm, hyperparameters, and seed. The elastic net stacker learns which models contribute unique signal. In a hackathon, building many simple models with diversity beats spending hours tuning one model perfectly.

Key Takeaway

In a 20-hour hackathon, the two highest-leverage decisions were feature selection strategy and ensemble diversity. Standard greedy selection would have missed interaction-dependent features. Standard single-model optimization would have capped the ceiling. The stochastic subset search and 14-model blend are what pushed the result to Rank 6. The lesson I carry forward: feature interaction spaces are too large for sequential search. Randomized exploration is not just a shortcut — it discovers combinations that greedy methods structurally cannot find.

Design insight: Feature interaction spaces are too large for greedy search. Stochastic sampling across many random subsets explores combinations that sequential selection structurally cannot find — and in time-constrained settings, it's the only viable path to the global optimum.

FAQ

What is the key takeaway from "STEG Fraud Detection"?

Feature interaction spaces are too large for greedy search. Stochastic sampling across many random subsets explores combinations that sequential selection structurally cannot find — and in time-constrained settings, it's the only viable path to the global optimum.

Who wrote this and what is it about?

This was written by Mahmoud Trigui, Senior Data Scientist. Detecting fraudulent meter manipulation for Tunisian utility STEG (200M Dinars lost). 15 years of billing history, 5.6% fraud rate. Stochastic feature selection across 100 LightGBM iterations, 14-model diversity ensemble with elastic net meta-learner. 5 encoding strategies for high-cardinality categoricals. Ranked 6th out of 191 at AI Hack Tunisia 2019.

Comments

Machine Learning Feature Engineering MLForecast Time Series Decomposition Forecasting LightGBM XGBoost Catboost Clustering Segmentation NLP LLMs Web App R Markdown SQL Oracle DB SAS-Guide SAS E-Miner Dataiku BigQuery GCP Python R CRISP-DM Hypothesis Testing ANOVA Data Analytics Dimensionality Reduction Recommendation System Network Analysis Geospace Analysis Spatial Data Embedding Sampling Techniques Decision Rules Data Storytelling CVM Churn Fraud Detection Sentiment Analysis Topic Modeling IBM Watson PowerBI Looker Studio VBA Statistical Learning Ensemble Modeling Stacking Cross-Validation Profiling ABT Construction Plumber Tidyverse Shiny Prophet Deep Learning Scikit-Learn JSON SAS Programming Git VS Code CSS Styling Automated Reporting Outlier Detection Temporal Clustering Startup Survival Pre-Valuation Modeling K-Means Decision Trees Data Science Predictive Modeling SVM LDA Text Classification Weight Prediction Pattern Recognition Real-Time Detection Community Detection Pipeline Automation Data Quality Checks Data Reliability Specification Mapping Business Strategy Marketing Campaigns Try & Buy Frameworks KPI Dashboards Network Quality Sales Analytics Mentoring Statistics Lecturer Remote Work Hybrid Work Consulting Contract Full-Time Freelance Sofrecom Orange Group Tunisia Telecom Kiota Intelligence VC Analytics Series A Prediction Pre-Valuation Modeling Production ML Applied AI Prompt Engineering Business Forecasting Decision Systems Graph Analytics Household Detection Multi-SIM Detection FTTH Forecasting Audit Extraction Infrastructure Classification Pydantic GPT-4 OpenAI API Base64 Classification Zindi Codementor LAAS-CNRS ESSAI MIT xPRO Tunisia ML Competition Cell Tower Analysis Uber Logistics Uber Cape Town Necessary Condition Analysis Behavioral Signals Spike Smoothing Observation Unit Design Dendrogram Ward Clustering VIF Target Encoding Machine Learning Feature Engineering MLForecast Time Series Decomposition Forecasting LightGBM XGBoost Catboost Clustering Segmentation NLP LLMs Web App R Markdown SQL Oracle DB SAS-Guide SAS E-Miner Dataiku BigQuery GCP Python R CRISP-DM Hypothesis Testing ANOVA Data Analytics Dimensionality Reduction Recommendation System Network Analysis Geospace Analysis Spatial Data Embedding Sampling Techniques Decision Rules Data Storytelling CVM Churn Fraud Detection Sentiment Analysis Topic Modeling IBM Watson PowerBI Looker Studio VBA Statistical Learning Ensemble Modeling Stacking Cross-Validation Profiling ABT Construction Plumber Tidyverse Shiny Prophet Deep Learning Scikit-Learn JSON SAS Programming Git VS Code CSS Styling Automated Reporting Outlier Detection Temporal Clustering Startup Survival Pre-Valuation Modeling K-Means Decision Trees Data Science Predictive Modeling SVM LDA Text Classification Weight Prediction Pattern Recognition Real-Time Detection Community Detection Pipeline Automation Data Quality Checks Data Reliability Specification Mapping Business Strategy Marketing Campaigns Try & Buy Frameworks KPI Dashboards Network Quality Sales Analytics Mentoring Statistics Lecturer Remote Work Hybrid Work Consulting Contract Full-Time Freelance Sofrecom Orange Group Tunisia Telecom Kiota Intelligence VC Analytics Series A Prediction Pre-Valuation Modeling Production ML Applied AI Prompt Engineering Business Forecasting Decision Systems Graph Analytics Household Detection Multi-SIM Detection FTTH Forecasting Audit Extraction Infrastructure Classification Pydantic GPT-4 OpenAI API Base64 Classification Zindi Codementor LAAS-CNRS ESSAI MIT xPRO Tunisia ML Competition Cell Tower Analysis Uber Logistics Uber Cape Town Necessary Condition Analysis Behavioral Signals Spike Smoothing Observation Unit Design Dendrogram Ward Clustering VIF Target Encoding