The disposable prototype that shipped to production
The prototype worked in the demo. Nobody scheduled the follow-up ticket to add authentication properly, because the demo was the deliverable and the deliverable shipped. Six weeks later it's a live URL with a database behind it, and nobody remembers it was ever meant to be disposable.
What "disposable" actually protects you from
A disposable prototype is supposed to be cheap because it gets thrown away. The entire cost model behind generating five variants and keeping one assumes the four you didn't keep, and the corners cut to make all five fast, never see a real user.
That assumption breaks the moment a stakeholder likes the prototype enough to just... use it. The corners that were fine for a throwaway are the same corners a production app can't survive with, and nothing about the tool that generated it flags which category it's currently in.
What that actually costs, measured
Security firm Escape scanned 5,600 live, deployed applications built with AI app-generation tools, most heavily Lovable, using passive testing that avoided disrupting the targets.
| Found across 5,600 live apps | Count |
|---|---|
| Vulnerabilities | 2,000+ |
| Exposed secrets (API keys, credentials, tokens) | 400+ |
| Instances of exposed personal data | 175 |
These are not toy repos. They are apps that were already running when the scan found them.
Amazon's own internal experience with the same category of tool tells a matching story at a different scale. On March 5, 2026, its AI coding agent Kiro was linked to a roughly six-hour outage estimated at 6.3 million lost orders. Amazon disputed a separate December 2025 incident's AI attribution as user error, which is itself worth taking seriously: not every failure blamed on an agent is actually the agent's fault. But the March incident, and the internal review that followed four Sev-1s in one week, are consistent with the same root pattern Escape found in the wild: unreviewed, agent-generated changes reaching production without the checks a human would have insisted on for anything permanent.
The corner that gets cut is always the invisible one
A missing button is obvious the moment you look at the screen. A missing authorization check is invisible right up until someone who shouldn't have access finds it. AI-generated code fails the same way design work fails silently: it renders fine, it demos fine, and the actual defect has no visual signature at all.
That is exactly why "it looked done" is not evidence it's safe to keep. Looking done was the one thing the disposable version was optimized for.
The five-minute question before anything goes live
Before a prototype crosses from "thing we're evaluating" to "thing a real user can reach," ask one question out loud in the room: has anyone actually looked at what happens if someone sends this a request it wasn't built to expect?
If the honest answer is no, that's not a reason to rebuild from scratch. It's a reason to run one focused security pass before the URL stops being a demo link and starts being a product. Fifteen minutes, pointed at the four categories the Escape scan actually kept finding:
Audit this codebase for the failure modes that show up in AI-generated apps
that shipped without review. Do not fix anything yet. Report findings only.
Check, in this order:
1. Authorization. For every route and every data read, can a signed-in user
reach another user's records by changing an ID? Is any check client-side
only?
2. Secrets. Any API key, token, connection string or credential committed to
the repo, bundled into client-side code, or exposed via a public env var?
3. Personal data. What PII does this store, which endpoints return it, and is
any of it reachable unauthenticated?
4. Input handling. Which endpoints accept input that reaches a database, a
file path, or a shell command without validation?
For each finding: the file and line, what an attacker sends, and what they
get back. Rank by what is reachable without logging in.
Point that at the prototype before it goes live, not after. The scan found what it found precisely because nobody ran anything like it while the app was still a demo.