Not All AI Assistance Deserves Equal Trust
The useful question is not whether to use AI, but what kind of work you are delegating and what happens if it is wrong. I sort AI assistance into four levels — and match the verification effort to the consequence.
TL;DR: Trust in AI assistance should scale with consequence, not with confidence in the output. Inspect code completion, test generated code, verify AI critique against evidence, and treat decision support as input for judgment rather than delegated judgment. The closer AI gets to framing the problem, choosing the observation unit, or designing the validation strategy, the stronger the verification has to be. AI does not replace data science judgment — it makes good judgment more valuable, because the scarce skill becomes knowing what to check.
Visual Summary
The Problem
The AI conversation in most data science teams is stuck on a yes-or-no question: should we use AI? That framing hides the only thing that actually matters. Assistance that fills in an import is not the same act as assistance that writes the validation split, and neither is the same as assistance that proposes how the prediction problem should be framed in the first place.
When trust is treated as binary, everything gets the same amount of it. Low-risk output gets reviewed carefully because it is small and visible, while the high-consequence output — a feature that quietly leaks, a validation design that does not match deployment, a framing that turns a ranking problem into a classification problem — sails through because it reads well and looks reasonable.
That is the failure mode I actually worry about. AI output is fluent by construction, so plausibility is not evidence. A suggestion that has never been checked behaves exactly like a finding that has, and nothing in the output itself tells the difference.
The Approach
So I stopped asking whether AI output can be trusted and started asking what kind of work I am delegating, and what the cost of being wrong would be. Four levels, ordered by how much of the project rides on the answer being right.
Level 1 — Code completion
Imports, function signatures, boilerplate, small syntax fixes. Fast to produce and fast to inspect, because the whole result fits on the screen and I know exactly what it should look like. I still review anything that touches data handling, security, dependencies, or production logic.
Level 2 — Code generation
An Optuna objective, a feature-engineering function, a time-series split utility, plotting or API scaffolding. This saves real implementation time, but the code is an untested claim: it needs tests, review, and validation against my own data before it goes anywhere near a pipeline I care about.
Level 3 — Review and critique
Possible leakage paths, missing validation checks, alternative feature ideas, weak assumptions in a pipeline. This is where AI is at its most leveraged, because it is a tireless second reader that can surface blind spots I stopped seeing. The catch is that its output is a hypothesis, not a finding — I verify every critique against data lineage, code and evidence.
Level 4 — Decision support
Framing the prediction problem, choosing the observation unit, comparing validation strategies, evaluating architecture trade-offs. Most useful and most dangerous at the same time, because a plausible answer here can shape the entire project before a single line of code is written.
The levels are not a ranking of tools. The same assistant shows up in all four, and the same task can move up or down depending on the consequence. What changes is the verification I attach to it.
My Practical Trust Policy
Once the levels are explicit, the policy writes itself: the verification requirement increases with consequence. Four rules, one per level, that I can apply without re-litigating the debate every time an assistant suggests something.
Complete → Inspect quickly
Level 1 output gets a fast, deliberate read. If it touches data, security, dependencies, or production logic, it gets a real review, not a glance.
Generate → Test it
Level 2 output earns its place through tests and validation against my data. Correctness is established by running it, not by reading it.
Critique → Verify with evidence
Level 3 suggestions are treated as hypotheses. I check each one against the code, the schema, the data lineage, and the metric before I accept it.
Support decisions → Input, not authority
Level 4 output is used to generate options and sharper questions. The decision stays mine, supported by domain context, data evidence, and experiments.
Outcome
What changed in practice is speed in the right places. On levels 1 and 2 I delegate more, because the cost of being wrong is bounded and the cost of writing boilerplate by hand is pure waste. On levels 3 and 4 I have become deliberately slower, because that is where an unverified assumption is expensive enough to reshape the whole project.
The second change is conversational. Instead of arguing about whether AI is allowed, the discussion moves to which level a task sits at and what verification it needs. That is a much better argument, because it can be settled with evidence instead of preference.
None of this makes me slower overall. It moves the effort from typing to checking, which is where the leverage was anyway. AI compresses implementation time; it does not compress accountability, and it cannot replace the judgment that decides which work is worth delegating at all.
Key Takeaway
Design insight: Trust in AI assistance should scale with consequence, not with confidence in the output. Inspect code completion, test generated code, verify AI critique against evidence, and treat decision support as input for judgment rather than delegated judgment. The closer AI gets to framing the problem, choosing the observation unit, or designing the validation strategy, the stronger the verification has to be. AI does not replace data science judgment — it makes good judgment more valuable, because the scarce skill becomes knowing what to check.
Related
AI-Assisted Coding Moved the Hard Part of Data Science to Specification →AI Can Review Feature Code. It Cannot Approve It →AI Pre-Flight Review Before You Write Model Code →Delegate the Optuna Scaffold. Not the Tuning Strategy. →Should This Be a Rule, a Model, or an LLM? →A Valid LLM Response Is Not Necessarily a Safe Decision →Claude Code Project Instructions: Layered CLAUDE.md →Open-Source AI Coding Agents: Which One Fits Your Workflow? →One Row Is a Modelling Decision: Customer or State? →AI Maturity Levels: From Demo to Platform →
FAQ
What are the four levels of AI assistance?
Level 1 is code completion: imports, function signatures, boilerplate and small syntax fixes, which are fast to inspect. Level 2 is code generation: an Optuna objective, a feature-engineering function, a time-series split utility or API scaffolding, which must be tested and validated against your own data. Level 3 is review and critique: leakage paths, missing validation checks, alternative feature ideas and weak assumptions, where the suggestions are hypotheses rather than findings. Level 4 is decision support: framing the prediction problem, choosing the observation unit, comparing validation strategies and weighing architecture trade-offs. The verification requirement increases with consequence.
How should AI-generated code be verified before you use it?
Read it, run it, and validate it against your own data. Generated code is an untested claim: a plausible implementation can still be wrong about your schema, your time semantics, or your split logic. Tests, data checks and a review pass are what establish correctness, and anything touching data handling, security, dependencies or production logic deserves a real review rather than a quick inspection.
Can AI review my model pipeline for data leakage?
It can help you look, but it cannot approve. Used at level 3 it is a tireless second reader that surfaces suspicious patterns: features that might use future information, splits that do not reflect deployment, missing signals, and assumptions that could make a result misleading. Every one of those suggestions is a hypothesis. Verifying data lineage, feature timing and prediction-time availability against the actual data stays a human responsibility.
Why is AI decision support the most dangerous level of assistance?
Because a plausible answer at that level can influence the whole project before a single line of code is written. Framing the problem, choosing the observation unit (customer, transaction or account) and designing the validation strategy decide what the model is even allowed to learn, and an error there stays invisible in every later metric. So use AI for options and sharper questions, and verify with domain context, data evidence, experiments and human judgment. A plausible recommendation is not a correct decision.