Maintaining AI-built apps
The invisibility problem
An agency audited 26 AI-built apps and found 14 recurring problems. The owner could have spotted one of them. The other 13 are a property of code, and you can design them out before launch.
7 min read
You built it in an afternoon. It works, mostly. You've been using it for a couple of weeks, your partner has, you sent the link to a friend who said "nice" and then never mentioned it again, which you took as a good sign.
Then six weeks in something breaks. What I want to talk about is why that bug is the least interesting thing about the situation, because in most of these apps the bug has been there since the first day, and nothing about the app could have told you.
#Thirteen of fourteen
The best data I've come across is from an agency that audits vibe-coded apps as their actual day job. They went through 26 of them at the code level and published a list of the 14 problems that kept turning up. I read the list trying to work out which of the 14 the owner could realistically have caught on their own, just by using the app. I got to one. The other 13 all needed something the owner didn't have handy: a second account, a stranger poking at it, real traffic, or a customer who got annoyed enough to write in.
Some of the numbers, from the 21 apps in that batch that weren't the agency's own work. 11 had a privileged endpoint that answered without a session. 7 allowed a confirmed cross-user action, so one customer could get at another customer's data. 13 had no rate limit on their most expensive endpoint, and 17 ran no error tracking whatsoever.
From the owner's chair, none of those four look like anything. The endpoint that forgot its session check behaves exactly like the one that remembered, as long as the person calling it is you. The missing rate limit costs nothing until traffic shows up, and it's mostly the owner's own traffic for a while. And with no error tracking, when a request eventually does fail, nobody is told, so the app looks fine.
A security firm did a much bigger scan and landed in the same place. They crawled 1,072 vibe-coded apps built on one popular hosted database. Over 300 had shipped the database key in the client-side JavaScript. 172 would let anyone delete data without logging in. Presumably every one of those had been used by whoever built it, and nothing had looked wrong.
#It's the code, not the model
It's tempting to read those numbers as a verdict on AI. I don't think that's what they are, though I'll admit I'd be suspicious of me saying that, given what I'm selling, so let me try to earn it.
The quality gap is real. CodeRabbit went through 470 open-source pull requests and found AI-co-authored changes carried about 1.7x as many issues as human-written ones, with security vulnerabilities up to 2.74x as common. In the 2025 Stack Overflow survey, 66% of developers said their biggest frustration was "AI solutions that are almost right, but not quite." The 2025 DORA report found AI adoption correlating with higher throughput and, still, with lower delivery stability. Those numbers will move as the models improve. Fine.
But notice what kind of problem every one of those is describing. Code that looked right and failed when it ran. Code that passed review and then fell over in production, with a team shipping more and getting paged more. In every case the fault was sitting there the whole time and nobody could see it until the code executed. The audit found the same thing. It just measured it with a different ruler.
Code has been like this since long before anyone pointed a model at it. A reference in a codebase might be a dynamic import, or a string glued together at run time, or a route that a framework registers by naming convention, or a permission check that exists on one path and is missing on the path right next to it. You can't read the code and know what depends on what, not reliably, not even if you wrote it. The only way to find out whether a change is safe is to run it under every condition that matters. And the conditions that matter are the exact ones from the audit list: someone who isn't you, real traffic, a second account.
That was all true in 2015 too. The difference now is volume. A lot more code is getting produced, much faster, by people who were never going to read it, and the same old invisibility applies to every line.
#Why every answer today is an audit
Search for what to do about a broken vibe-coded app and every result goes the same way. Book an audit. A human reads the codebase, finds the 13 invisible problems, fixes them or tells you to start over. A couple of the hosting platforms offer a rollback instead, which gets you back to how things were two weeks ago.
Neither is a bad answer. Both happen after the fact, though, and cost what a human costs, and the question they answer is "what broke?" when the question I actually had, every time, was "what would break if I did this?"
I looked for a while and couldn't find anyone in those results asking whether the app could have been checked before it went live.
#What it looks like when an app can be checked before it runs
That's the question Tessryx exists to answer, and the whole thing hinges on one decision I made early: the agent doesn't write code. It writes definitions.
A page, an API route, a datafile, a schema, a scheduled job, a call out to some other service. On Tessryx each of those is a definition with its own schema rather than a pile of code, and each one declares what it references. I keep coming back to that last part because it's the whole trick. It means the platform can know things in advance that a codebase can only find out by running.
When the agent saves a definition, the platform checks it against its rules right then, so a broken one never gets as far as running. Before publishing, the agent can preview a page as a specific signed-in viewer, which is the second-account check from the audit list without needing a second account, and it can do the same as a stranger. The references are declared, so the platform can walk them in either direction. That's the day-two question, what breaks if I change this, answered directly rather than by running things and watching.
Here's that last one, live, rather than described. This is the real dependency graph of my birthday invite site (magic-link invites, an admin behind sign-in, nothing fancy), exactly as the platform's own analyzer reports it. Hover a piece.
- custom domain
- endpoint
- workflow
- datafile
- schema
$ analyze_resource --kind workflow --resource birthday/theme --hypothetical delete
deployable: false · errors: 7
- errorworkflow:birthday/home-render would break: steps[0](theme).sub_workflow.workflow references "birthday/theme@published", which would no longer exist
- errorworkflow:birthday/invite-render would break: steps[0](theme).sub_workflow.workflow references "birthday/theme@published", which would no longer exist
- errorworkflow:birthday/admin-render would break: steps[0](theme).sub_workflow.workflow references "birthday/theme@published", which would no longer exist
- errordynamic_endpoint:birthday/home would break: depends on birthday/home-render, which would break
- errordynamic_endpoint:birthday/i/:token would break: depends on birthday/invite-render, which would break
- errordynamic_endpoint:birthday/admin would break: depends on birthday/admin-render, which would break
- errorcustom_domain:birthday.nickpage.tech would break: depends on birthday/home, which would break
The red chain is the obvious thing to look at. Pull out the shared theme workflow and the analyzer reports three page workflows broken, then the three endpoints that run them, then the custom domain whose root points at one of those endpoints. Seven findings from one hypothetical delete, with nothing actually deleted, and nobody had to read the app to get them.
The amber is the part I actually like. Hover a datafile and the analyzer reports no static referrers, then flags four workflows whose load target is computed at run time and therefore can't be checked statically. It could have stayed quiet about those. It doesn't, and that's the bit I'd want from any tool claiming to audit anything: I want to know where it can't see.
#Where this stops helping
Definitions don't make an app correct. An agent can still build a page that shows the wrong thing, or a workflow with a logic error in it, and the analyzer will report that everything resolves, because it does.
The kind of problem that gets to hide is different, though. Missing auth and cross-user access get caught before publish, or at least get asked about. So does a reference that quietly points at nothing. None of those wait around for a customer to complain. The 13 of 14 shrinks to the problems that were always going to need a human, which is where I'd rather the humans spent their time anyway.
#What to do this week
If you've got an AI-built app in production right now, the checklists in the sources below are good and they mostly agree on where to start. Rotate any key that has ever shipped to a browser, and search the built JavaScript for it, not just your source. Then go through every privileged route and confirm it actually checks for a session, which is tedious and is also where 11 of the 21 apps above failed. Turn on error tracking, because otherwise your customers are your error tracking. Rate-limit the expensive endpoint. Call it an afternoon.
If you're about to build one, decide where the checks are going to happen before you decide which model writes it. A site or app made of definitions gets the dependency question answered before it's previewed and published. When something does go wrong there's a trace of the run to read. And when you hand it to someone else, the structure comes with a lock on it. You bring whatever agent you already use. Connecting it takes a few minutes.
#Sources
- Axonbuild, The biggest problem with vibe coding, from 26 audits
- Symbiotic Security, We scanned 1,072 vibe-coded apps: 98% had security flaws
- CodeRabbit, State of AI vs human code generation report
- Stack Overflow, 2025 Developer Survey: AI
- Google Cloud, Announcing the 2025 DORA report