Maintaining AI-built apps
Why the answer to AI tech debt is an audit, and why it shouldn't have to be
The evidence that AI code accumulates debt faster is solid. Every remedy on offer is a human reading it afterwards. That is a property of code, not of AI.
5 min read
"AI tech debt" is a contested phrase. Half the people using it mean "my team ships faster now and I'm nervous about it." The other half mean something specific and measurable, and the measurements have got good enough over the last year that it's worth being precise about which one you're talking about.
This post is about the specific one. The way code written by an agent accumulates problems the person who owns it can't see, and why every remedy currently on offer amounts to hiring a human to read it afterwards.
#The evidence, briefly
Four independent measurements. None of them come from vendors selling a fix, which I mention because I am one.
Churn and duplication. GitClear's analysis of 211 million lines of code found churn, meaning lines rewritten shortly after being written, rising from a pre-AI baseline of 3.3% to 7.1%. Over the same period copy-pasted lines went up, and "moved" lines, which are the signature of refactoring, fell to under 10% of changes. Blocks of five or more duplicated lines increased eightfold in a single year. If you drew technical debt on a chart it would look like that. More code with less structure, and a lot of it rewritten soon after it landed.
Defect rate. CodeRabbit's review of 470 open-source pull requests found AI-co-authored changes carrying about 1.7 times as many issues as human ones, with security vulnerabilities up to 2.74 times as frequent and error-handling gaps nearly twice as common.
Stability. The 2025 DORA report found AI adoption now correlates with higher delivery throughput and, still, with lower delivery stability. Teams ship more, and more of it breaks.
The perception gap. This is the one I keep thinking about. METR ran a randomized trial where experienced developers worked real issues in repositories they knew well, with and without AI tools. With the tools they took 19% longer, while believing they'd been about 20% faster. The speed number gets all the attention. The finding that matters for debt is that the people closest to the code misjudged what it was costing them by nearly forty points.
So, taken together: AI-written code has more defects and more duplication and more churn than human-written code, it makes delivery less stable, and the people responsible for it underestimate all of that.
#Why the only remedy is a person reading it
Search for what to do about it and the market has one answer, which is an audit. A senior engineer reads the codebase, finds what the agent got wrong, and either fixes it or tells you to rebuild. Some of these are excellent. They're all the same shape though.
They have to be the same shape, because of what code is. You can't ask a codebase what it depends on. A reference might be a dynamic import, or a string assembled at run time, or a permission check that exists on one route and not its twin. The only way to find out whether the app is sound is to run it in every condition that matters, or to have somebody who understands it read the whole thing and imagine those conditions. Rollbacks are the other remedy you'll see offered, and they amount to the same admission. The platform can't tell you what a change will do, so it offers to undo it once you've found out the hard way.
An audit is remediation. By the time you're paying for one, the debt has already been taken on. The question I'd rather ask is whether there's a way to build where the invisible kind of debt can't accumulate in the first place. That's a question about what the agent writes into, more than about which model does the writing.
#Prevention is a property of what the agent writes
Tessryx makes one design decision that moves the debt into a different category. The agent doesn't produce code. It produces definitions. A page, an API route, a datafile, a schema, a scheduled job, a call to an outside service, each a bounded declarative definition with its own schema and its own declared references.
Definitions can carry debt too, so let me be exact about what changes and what doesn't.
Debt that can't be taken on. A definition is validated against its rules when it's saved. So the categories audits find most often, a reference to something that doesn't exist, a route wired to a workflow that got renamed, a required field that was dropped, get refused at save time rather than discovered in production. Every reference is declared, which means the platform can walk them in both directions and answer "what breaks if I change this?" before anything is changed, including a hypothetical delete that reports its findings without deleting. Where a reference genuinely can't be resolved statically, because the target is computed at run time, the analysis says so instead of going quiet. The live version of all that is in The invisibility problem, running against a real app of mine.
Churn, but bounded. A published definition is a frozen version. Changing it means a new version, previewed before it's published, with the old one still serving until that moment and available to publish again. Churn still happens. But each change is visible and dated and can be reversed, and "what was it like before" is a dropdown rather than an archaeology project. Every run is traced, so when something fails there's a record of what ran and with what instead of an absence.
What's still yours to carry. An agent can still build the wrong thing. A workflow can have a logic error that resolves perfectly and produces the wrong number. A definition can't make a product decision good, and nothing here claims to. What it does is move the invisible categories, the ones that needed a stranger or a second account or a customer complaint to surface, over into the visible ones, where the person who owns the app can see them without paying somebody to look.
#What this means if you're holding the debt now
If you have an AI-built codebase in production, get the audit. The evidence above is the reason to, and a good one will show you the same handful of things: routes without session checks, keys that shipped to a browser, missing error tracking, no rate limit on the expensive endpoint. Fix those first.
Then decide where the next one gets built. If it needs custom code, it'll carry code debt, and you should plan to audit it again in a year. If it's a site, a portal, a form and a table, a page over an API, or a scheduled job, which covers most of what people are actually shipping, it can be built out of definitions instead. The kind of debt that needs an audit to find then doesn't get taken on in the first place, and the audit money can go on something else.