The parts of a system you should never let AI decide

Directing AI through a build doesn't mean stepping back from every decision. It means being precise about which decisions are yours, always, no matter how good the model gets.

Most of the discourse about AI-built software treats "who decided this" as a single question with a single answer: either a human wrote it or the model did. In practice, building three real systems this way — a research tool, a feedback analyzer, a lead pipeline — taught me that the question has to be asked at the level of individual decisions, not the project as a whole. Some of those decisions I handed to the model without a second thought. Others I kept, deliberately, every time.

Here's the actual list, with the reasoning behind each one.

What counts as a failure

A model will happily paper over a gap. Ask it to classify a piece of text and it will pick the closest available label even when nothing actually fits — because guessing looks more finished than admitting uncertainty. That instinct is exactly backwards for a system whose entire value proposition is trustworthiness.

In one of my tools, feedback that doesn't match any known category is labeled uncategorized rather than forced into the nearest bucket. That's not a technical limitation — it would have been trivial to make the model always pick something. It's a decision about what honesty looks like in the output, and it's not one I was willing to let get optimized away in the name of a cleaner-looking result.

How a failure should behave

Separately from what counts as a failure: what happens when one occurs? A lead-processing pipeline I built rejects malformed input at the first stage — a missing email, an empty message — and stops there, visibly, with a specific reason attached to the record. It would have been just as easy to let a bad record limp through the rest of the pipeline collecting default values, and just as easy to make it disappear silently rather than showing up in a queue as "rejected." Neither of those is a coding problem. They're both judgment calls about whether the system should fail loud or fail quiet, and I don't think that's a call you hand off.

What the system is allowed to claim

This is the one I'd put first if I had to rank them. Every one of these tools makes claims — a research brief says a finding is supported by a source, a scoring model says a lead earned 84 points, a customer-intelligence report says a theme is trending negative. The rule I set, and checked constantly, is that every claim has to trace back to something real: a quote, a data point, a rule that fired. Nothing gets to sound more confident than the evidence underneath it.

This matters more with AI in the loop, not less, because a model is a genuinely excellent tool for making an unsupported claim sound completely reasonable. That's a feature when you want persuasive writing. It's a liability when the whole point of the tool is to be trusted.

Where the scope boundary sits

What a fixed-scope system does and doesn't do is a business decision wearing a technical costume. "Should this also handle CSV uploads" or "should the scoring model factor in company size" are the kind of questions that are easy to say yes to in the moment and expensive to have said yes to six months later. Every time I've let scope drift because a feature seemed easy to add, it's been a mistake I noticed later, not one I noticed at the time — which is exactly why it has to be a standing decision, made in advance, not a per-request judgment call.

What "done" means

Last, and it sounds almost too simple to state: deciding when something is finished. A model doesn't have a stopping condition of its own — it will keep refining, adding, and elaborating for as long as you let it, and every one of those passes looks like progress. Knowing that the fourth version was actually worse than the second, or that a feature nobody asked for doesn't belong in the first release, is not something I've found a way to delegate. It's the whole job, if I'm honest about it.

None of this is an argument against directing AI through a build. It's the opposite — it's what makes doing it responsibly possible. The parts on this list are exactly the parts that don't get easier or safer to hand off as the tools improve, because they were never really implementation questions to begin with.