What actually happens when you have AI build a real backend
Not a hot take. A log of the specific, unglamorous problems that came up building three working systems, and what actually catching them looked like.
The honest version of "I used AI to build a backend" isn't a straight line from prompt to working software. It's closer to directing a very fast, very literal collaborator who will confidently produce something that runs — and running is not the same thing as correct. Here's what that actually looked like across three projects: a research-clustering tool, a feedback-analysis pipeline, and a lead-scoring engine.
The bug that only showed up under real load
The first version of the document-clustering logic worked perfectly in testing — with two or three sample documents. The moment I ran it against a realistic batch, nearly everything landed in its own isolated cluster instead of grouping with related material. The clustering threshold had been tuned against a toy dataset and simply didn't hold at real scale. The fix was a one-line change to a distance parameter. Finding it required actually running the thing against realistic data and noticing the output looked wrong, which is a step that's easy to skip when the code "works" in the sense of not throwing an error.
The variable that was set and never used
A hover effect across three product previews was supposed to shift to a different accent color per product — orange, green, blue. All three shifted to the same orange, every time. The CSS referenced a variable that was never actually assigned anywhere; it just silently fell back to a default, and the fallback happened to look plausible enough that nobody would notice without deliberately checking each one side by side. That's the category of bug that doesn't announce itself. It just quietly ships.
The label that didn't fit its own column
A two-column layout — a short label next to a longer explanation — worked fine when the label said "Problem." The same layout broke visibly when a longer label, "Contribution," didn't fit the column width and bled into the text next to it. The column had been sized for the first label that happened to get typed in, not for the actual range of labels the layout needed to support. It's a small thing, and it's also exactly the kind of small thing that only surfaces when someone actually looks at the rendered page rather than the code that produced it.
The fix that got built and then never deployed
Maybe the most instructive one: a genuine, verified fix — reordering a navigation menu so visitors saw context before data — sat finished and correct for several turns while the live site kept showing the old broken order. The fix wasn't wrong. It just hadn't actually gone out yet, and nothing about "the code is correct" tells you whether "the code is live" is also true. Those are different questions, and conflating them is an easy way to think something is solved when it isn't.
What all four of these have in common
None of them were caused by the AI writing bad code in the sense of syntax errors or broken logic. Every one of them ran. Every one of them looked done. What made each one a real bug was a gap between what the code did and what it was supposed to do — a gap that's invisible from the code itself and only visible from the output, the rendered page, the deployed site, the edge case nobody thought to test.
That's the actual shape of directing AI through a real build: not writing every line, but being the one who runs it, reads the output with real skepticism, and keeps looking until the gap between "runs" and "correct" closes. The tools are good enough now that the code showing up fast was never the hard part. Noticing what's quietly wrong with code that looks completely fine — that's still the job.