Not All AI Assistance Deserves Equal Trust

The useful question is not whether to use AI, but what kind of work you are delegating and what happens if it is wrong. I sort AI assistance into four levels — and match the verification effort to the consequence.

4 Assistance Levels Verification by Consequence AI Code Review Human Judgment

TL;DR: Trust in AI assistance should scale with consequence, not with confidence in the output. Inspect code completion, test generated code, verify AI critique against evidence, and treat decision support as input for judgment rather than delegated judgment. The closer AI gets to framing the problem, choosing the observation unit, or designing the validation strategy, the stronger the verification has to be. AI does not replace data science judgment — it makes good judgment more valuable, because the scarce skill becomes knowing what to check.

Visual Summary

The Problem

The AI conversation in most data science teams is stuck on a yes-or-no question: should we use AI? That framing hides the only thing that actually matters. Assistance that fills in an import is not the same act as assistance that writes the validation split, and neither is the same as assistance that proposes how the prediction problem should be framed in the first place.

When trust is treated as binary, everything gets the same amount of it. Low-risk output gets reviewed carefully because it is small and visible, while the high-consequence output — a feature that quietly leaks, a validation design that does not match deployment, a framing that turns a ranking problem into a classification problem — sails through because it reads well and looks reasonable.

That is the failure mode I actually worry about. AI output is fluent by construction, so plausibility is not evidence. A suggestion that has never been checked behaves exactly like a finding that has, and nothing in the output itself tells the difference.

The Approach

So I stopped asking whether AI output can be trusted and started asking what kind of work I am delegating, and what the cost of being wrong would be. Four levels, ordered by how much of the project rides on the answer being right.

Level 1 — Code completion

Imports, function signatures, boilerplate, small syntax fixes. Fast to produce and fast to inspect, because the whole result fits on the screen and I know exactly what it should look like. I still review anything that touches data handling, security, dependencies, or production logic.

Level 2 — Code generation

An Optuna objective, a feature-engineering function, a time-series split utility, plotting or API scaffolding. This saves real implementation time, but the code is an untested claim: it needs tests, review, and validation against my own data before it goes anywhere near a pipeline I care about.

Level 3 — Review and critique

Possible leakage paths, missing validation checks, alternative feature ideas, weak assumptions in a pipeline. This is where AI is at its most leveraged, because it is a tireless second reader that can surface blind spots I stopped seeing. The catch is that its output is a hypothesis, not a finding — I verify every critique against data lineage, code and evidence.

Level 4 — Decision support

Framing the prediction problem, choosing the observation unit, comparing validation strategies, evaluating architecture trade-offs. Most useful and most dangerous at the same time, because a plausible answer here can shape the entire project before a single line of code is written.

The levels are not a ranking of tools. The same assistant shows up in all four, and the same task can move up or down depending on the consequence. What changes is the verification I attach to it.

My Practical Trust Policy

Once the levels are explicit, the policy writes itself: the verification requirement increases with consequence. Four rules, one per level, that I can apply without re-litigating the debate every time an assistant suggests something.

Complete → Inspect quickly

Level 1 output gets a fast, deliberate read. If it touches data, security, dependencies, or production logic, it gets a real review, not a glance.

Generate → Test it

Level 2 output earns its place through tests and validation against my data. Correctness is established by running it, not by reading it.

Critique → Verify with evidence

Level 3 suggestions are treated as hypotheses. I check each one against the code, the schema, the data lineage, and the metric before I accept it.

Support decisions → Input, not authority

Level 4 output is used to generate options and sharper questions. The decision stays mine, supported by domain context, data evidence, and experiments.

Outcome

What changed in practice is speed in the right places. On levels 1 and 2 I delegate more, because the cost of being wrong is bounded and the cost of writing boilerplate by hand is pure waste. On levels 3 and 4 I have become deliberately slower, because that is where an unverified assumption is expensive enough to reshape the whole project.

The second change is conversational. Instead of arguing about whether AI is allowed, the discussion moves to which level a task sits at and what verification it needs. That is a much better argument, because it can be settled with evidence instead of preference.

None of this makes me slower overall. It moves the effort from typing to checking, which is where the leverage was anyway. AI compresses implementation time; it does not compress accountability, and it cannot replace the judgment that decides which work is worth delegating at all.

Key Takeaway

Design insight: Trust in AI assistance should scale with consequence, not with confidence in the output. Inspect code completion, test generated code, verify AI critique against evidence, and treat decision support as input for judgment rather than delegated judgment. The closer AI gets to framing the problem, choosing the observation unit, or designing the validation strategy, the stronger the verification has to be. AI does not replace data science judgment — it makes good judgment more valuable, because the scarce skill becomes knowing what to check.

FAQ

What are the four levels of AI assistance?

Level 1 is code completion: imports, function signatures, boilerplate and small syntax fixes, which are fast to inspect. Level 2 is code generation: an Optuna objective, a feature-engineering function, a time-series split utility or API scaffolding, which must be tested and validated against your own data. Level 3 is review and critique: leakage paths, missing validation checks, alternative feature ideas and weak assumptions, where the suggestions are hypotheses rather than findings. Level 4 is decision support: framing the prediction problem, choosing the observation unit, comparing validation strategies and weighing architecture trade-offs. The verification requirement increases with consequence.

How should AI-generated code be verified before you use it?

Read it, run it, and validate it against your own data. Generated code is an untested claim: a plausible implementation can still be wrong about your schema, your time semantics, or your split logic. Tests, data checks and a review pass are what establish correctness, and anything touching data handling, security, dependencies or production logic deserves a real review rather than a quick inspection.

Can AI review my model pipeline for data leakage?

It can help you look, but it cannot approve. Used at level 3 it is a tireless second reader that surfaces suspicious patterns: features that might use future information, splits that do not reflect deployment, missing signals, and assumptions that could make a result misleading. Every one of those suggestions is a hypothesis. Verifying data lineage, feature timing and prediction-time availability against the actual data stays a human responsibility.

Why is AI decision support the most dangerous level of assistance?

Because a plausible answer at that level can influence the whole project before a single line of code is written. Framing the problem, choosing the observation unit (customer, transaction or account) and designing the validation strategy decide what the model is even allowed to learn, and an error there stays invisible in every later metric. So use AI for options and sharper questions, and verify with domain context, data evidence, experiments and human judgment. A plausible recommendation is not a correct decision.

Comments

Machine Learning Feature Engineering MLForecast Time Series Decomposition Forecasting LightGBM XGBoost Catboost Clustering Segmentation NLP LLMs Web App R Markdown SQL Oracle DB SAS-Guide SAS E-Miner Dataiku BigQuery GCP Python R CRISP-DM Hypothesis Testing ANOVA Data Analytics Dimensionality Reduction Recommendation System Network Analysis Geospace Analysis Spatial Data Embedding Sampling Techniques Decision Rules Data Storytelling CVM Churn Fraud Detection Sentiment Analysis Topic Modeling IBM Watson PowerBI Looker Studio VBA Statistical Learning Ensemble Modeling Stacking Cross-Validation Profiling ABT Construction Plumber Tidyverse Shiny Prophet Deep Learning Scikit-Learn JSON SAS Programming Git VS Code CSS Styling Automated Reporting Outlier Detection Temporal Clustering Startup Survival Pre-Valuation Modeling K-Means Decision Trees Data Science Predictive Modeling SVM LDA Text Classification Weight Prediction Pattern Recognition Real-Time Detection Community Detection Pipeline Automation Data Quality Checks Data Reliability Specification Mapping Business Strategy Marketing Campaigns Try & Buy Frameworks KPI Dashboards Network Quality Sales Analytics Mentoring Statistics Lecturer Remote Work Hybrid Work Consulting Contract Full-Time Freelance Sofrecom Orange Group Tunisia Telecom Kiota Intelligence VC Analytics Series A Prediction Production ML Applied AI Prompt Engineering Business Forecasting Decision Systems Graph Analytics Household Detection Multi-SIM Detection FTTH Forecasting Audit Extraction Infrastructure Classification Pydantic GPT-4 OpenAI API Base64 Classification Zindi Codementor LAAS-CNRS ESSAI MIT xPRO Tunisia ML Competition Cell Tower Analysis Uber Logistics Uber Cape Town Necessary Condition Analysis Behavioral Signals Spike Smoothing Observation Unit Design Dendrogram Ward Clustering VIF Target Encoding Coding Agents Open Source AI Aider Cline Continue OpenHands Goose Try & Buy Frameworks Machine Learning Feature Engineering MLForecast Time Series Decomposition Forecasting LightGBM XGBoost Catboost Clustering Segmentation NLP LLMs Web App R Markdown SQL Oracle DB SAS-Guide SAS E-Miner Dataiku BigQuery GCP Python R CRISP-DM Hypothesis Testing ANOVA Data Analytics Dimensionality Reduction Recommendation System Network Analysis Geospace Analysis Spatial Data Embedding Sampling Techniques Decision Rules Data Storytelling CVM Churn Fraud Detection Sentiment Analysis Topic Modeling IBM Watson PowerBI Looker Studio VBA Statistical Learning Ensemble Modeling Stacking Cross-Validation Profiling ABT Construction Plumber Tidyverse Shiny Prophet Deep Learning Scikit-Learn JSON SAS Programming Git VS Code CSS Styling Automated Reporting Outlier Detection Temporal Clustering Startup Survival Pre-Valuation Modeling K-Means Decision Trees Data Science Predictive Modeling SVM LDA Text Classification Weight Prediction Pattern Recognition Real-Time Detection Community Detection Pipeline Automation Data Quality Checks Data Reliability Specification Mapping Business Strategy Marketing Campaigns Try & Buy Frameworks KPI Dashboards Network Quality Sales Analytics Mentoring Statistics Lecturer Remote Work Hybrid Work Consulting Contract Full-Time Freelance Sofrecom Orange Group Tunisia Telecom Kiota Intelligence VC Analytics Series A Prediction Production ML Applied AI Prompt Engineering Business Forecasting Decision Systems Graph Analytics Household Detection Multi-SIM Detection FTTH Forecasting Audit Extraction Infrastructure Classification Pydantic GPT-4 OpenAI API Base64 Classification Zindi Codementor LAAS-CNRS ESSAI MIT xPRO Tunisia ML Competition Cell Tower Analysis Uber Logistics Uber Cape Town Necessary Condition Analysis Behavioral Signals Spike Smoothing Observation Unit Design Dendrogram Ward Clustering VIF Target Encoding Coding Agents Open Source AI Aider Cline Continue OpenHands Goose Try & Buy Frameworks