Before I write serious model code, I use an AI assistant to challenge five assumptions — problem framing, features, validation, leakage, and the baseline. It does not build the project; it makes the plan reviewable.
TL;DR: Use AI early to challenge assumptions, not late to justify results. An AI assistant is a useful second set of eyes on problem framing, feature candidates, validation design and leakage risk — but the business context decides the problem, the data-generating process decides the split, and the human expert verifies availability, timing and data lineage. The value is not AI-generated answers; it is reaching the modelling stage with assumptions that are visible, challengeable, and easy to test.
There is a familiar order of operations in a modelling project: load the data, build features, train something, tune it, and only then ask whether the problem was framed correctly. It is efficient to execute and expensive to unwind. A model can be perfectly well built around the wrong decision, validated on a split that does not match deployment, and reported with a metric that nobody in the business actually cares about.
I started using an AI assistant earlier in that process — not to write the pipeline, but to challenge my first assumptions before I invest hours in it. The assistant is a fast, tireless second reader. What it is not is a decision maker. The value is not in the answers it produces; it is in reaching the modelling stage with assumptions that are written down, visible, and challengeable.
My modelling checklist has five gates, and I run them before any serious model code. The assistant proposes; I verify. Together they are the validation strategy for the plan itself, not just for the model.
Start from the decision: what decision are we improving? Then check whether this is even a classification problem. It might be ranking, survival analysis, forecasting, anomaly detection, or a rules problem that no model should be solving at all. The assistant is good at surfacing alternatives; the business context is what decides between them.
Given the available schema, ask which categories of feature might be missing — behavioural changes, historical aggregates, relative-position features, calendar effects, entity relationships, external signals. Then do the part the model cannot: check availability, feasibility, meaning, and point-in-time correctness — does the feature exist at the moment the prediction is made? Suggested does not mean available, valid, or useful.
Ask what a realistic evaluation environment looks like, then let the data-generating process decide. Time-ordered data wants a rolling-origin backtest or a future holdout. Grouped entities want a group-aware split so the same customer cannot appear on both sides. Independent observations can take a random split. And a gap is sometimes required, when labels overlap or data arrives late.
I ask it to inspect feature logic and raise questions: could this use information unavailable at prediction time, is a target-derived aggregate fitted safely, does a rolling feature look into the future? It is genuinely useful at surfacing suspicious patterns, and it turns data leakage from an invisible risk into a named one. It is not proof of safety. A plausible answer is not a guarantee, and the verification of data lineage and timing stays with the human expert.
Before modelling, write down the simple baseline model to beat, the metric that actually matters at the relevant horizon, the operational constraints (latency, cost, availability, ownership), and the minimum improvement worth shipping. Metric selection is a modelling decision, not a reporting detail. Without those four, a model can improve a score that changes no decision. An assistant can help structure the criteria; it cannot judge whether the expected result is credible.
The practical effect is that arguments happen earlier and cheaper. A framing dispute that would have surfaced three weeks later, after a pipeline, a tuning sweep and a reporting slide, now happens on a whiteboard before any of that exists. The five gates are also reusable: they are the same review I apply to a colleague's design, so AI-assisted development of this kind does not skip the human step, it preps it. In an MLOps workflow it is cheap to run and expensive to omit, because it changes what model selection ever gets to compare.
The failure mode to watch is treating plausible output as a review result. The assistant widens the option set; it does not confirm availability, timing, or business meaning. Every suggestion still has to be checked against the schema, the data lineage, and the decision it is meant to support.
The same assistant asked to review a finished model tends to produce a polished confirmation of what already exists. Asked before the pipeline, it produces questions worth answering. The timing is the whole difference.
This is not a case for letting AI run the project. It is a case for spending ten minutes early on the five questions I would otherwise skip. Start coding with a clearer plan, then go and test every assumption in it.
Design insight: Use AI early to challenge assumptions, not late to justify results. An AI assistant is a useful second set of eyes on problem framing, feature candidates, validation design and leakage risk — but the business context decides the problem, the data-generating process decides the split, and the human expert verifies availability, timing and data lineage. The value is not AI-generated answers; it is reaching the modelling stage with assumptions that are visible, challengeable, and easy to test.
AI Can Review Feature Code. It Cannot Approve It → Not every analytics question is about what drives the outcome → A predictive model is not a decision system → The most dangerous label in ML is the one that looks correct but isn't! → Seed Sensitivity Is Part of Model Selection → Foundation Models Raise the Baseline → Encoding is a modeling decision, not a preprocessing checkbox → When the Label Does Not Exist: Define the Behaviour First → A Strong AutoML Baseline Can Beat Hand-Tuned Models → Automate Maintenance, Keep Judgment Human →
Use AI early to challenge assumptions, not late to justify results. An AI assistant is a useful second set of eyes on problem framing, feature candidates, validation design and leakage risk — but the business context decides the problem, the data-generating process decides the split, and the human expert verifies availability, timing and data lineage. The value is not AI-generated answers; it is reaching the modelling stage with assumptions that are visible, challengeable, and easy to test.
This was written by Mahmoud Trigui, Senior Data Scientist. AI as a pre-flight review layer: challenge problem framing, features, validation, leakage and the baseline before you build the pipeline.
No. It widens the option set and surfaces questions worth asking, but it does not decide. The business context chooses the problem, the data-generating process chooses the split, and a human expert verifies availability, point-in-time correctness and data lineage. Treat its output as suggestions to verify, never as a review result.
Five things: what decision the model is meant to support, and whether it is even a modelling problem; which feature categories are missing, and whether each one exists at prediction time; what a realistic validation environment looks like, such as a rolling-origin backtest, a group-aware split, or a random split; where data leakage could enter; and the baseline model plus the success criteria to beat. Answering these first turns hours of rework into a short design review.