The Demo Is Not the Product.
Eighteen checks for auditing a vibe-coded app across repository truth, user journeys, auth, data, UX, accessibility, performance, security, commercial assumptions, and regression control.

- 18
- 03
- Expected → Actual
- Decision-ready
Find the evidence gap. Open one focused check.
Search by risk, workflow, system boundary, or decision. Every field check keeps the same evidence-first structure, so you can compare findings without inventing certainty.
Eighteen checks. Three layers of product truth.
Open a check for its scenario, evidence lab, field test, role guidance, checklist, bounded prompt, verification questions, and source trail.

Scope and authority
An audit becomes unreliable when nobody can state what is being tested, which environment is in scope, what evidence is allowed, or what the auditor is forbidden to change.

Repo, release and environment identity
A passing test is meaningful only when it is tied to a specific code state, deployment, runtime, and data environment.

Product map and screen jobs
A product audit misses risk when it starts from the screens the auditor happens to notice instead of an inventory of routes, roles, actions, and data flows.

Accounts, auth, roles and permissions
Authentication proves who someone is. Authorization proves what that identity may read, change, export, share, administer, or recover.

Data persistence and migrations
A success toast is not persistence proof. Data is trustworthy only when the intended write survives validation, concurrency, reload, later reads, migrations, and failure recovery.

Feature state and completeness
Product status drifts when code existence, roadmap labels, help content, flags, screenshots, and live behavior are treated as the same kind of evidence.

First run and activation
A product can be feature-rich and still fail first use if access, setup, vocabulary, permissions, imports, or empty states prevent a new user from reaching a useful outcome.

Critical end-to-end journeys
Individual screens can pass while the journey between them fails through lost context, stale data, broken handoffs, role boundaries, or missing completion feedback.

States, feedback and recovery
Real products spend time loading, empty, partial, stale, disabled, denied, offline, invalid, retrying, and recovering. Those states are the product too.

Accessibility and responsive context
Responsive and accessible design succeeds when the same task remains understandable and operable under different constraints, not when the desktop screenshot merely shrinks.

APIs, integrations and contracts
Many production defects live at boundaries: client versus server validation, webhook delivery, third-party limits, retry semantics, version drift, and assumptions about what another service returns.

Observability, incidents and support
When a product fails, teams need enough correlated evidence to answer what happened, who was affected, where the request went, and what changed—without logging sensitive data.

Database and query health
A missing index, N+1 pattern, or connection issue becomes a performance finding only when evidence connects it to the workload and the observed cost.

Network, caching and payloads
Latency often comes from request waterfalls, oversized payloads, duplicate fetching, missing caching boundaries, or unnecessary third-party work rather than raw network speed.

Runtime, main thread and memory
Main-thread blocking and memory growth have different causes: expensive synchronous work delays interaction, while retained objects, listeners, timers, subscriptions, or caches can accumulate across time.

Security, privacy and isolation
Security review is not a generic vulnerability scan. It is verification that identity, authorization, input, secrets, storage, isolation, logging, and data lifecycle match the product’s real threat and privacy boundaries.

Commercial assumptions and retention
A technically healthy product can still fail if the customer is vague, the current workaround is good enough, switching cost is underestimated, onboarding is expensive, or repeat use never becomes a habit.

Regression, release and change control
A correction is not durable until the release process can detect the old failure returning and the team knows how to stop or roll back a bad change.
How to scope, investigate, report, and protect the result.The operating model, quick audit, master investigation prompt, finding schema, coverage matrix, FAQs, and related guides remain available in reliable HTML.
Why this guide exists
A vibe-coded product can look finished while the important parts remain unverified.
A polished screen does not prove the right code is deployed. A successful login does not prove authorization. A green toast does not prove persistence. A passing build does not prove the product journey. A fast demo does not prove the system will stay fast, secure, supportable, or worth paying for.
This guide treats an audit as a controlled investigation:
Scope → Identify → Map → Exercise → Measure → Explain → Prioritize → Verify → Protect
The goal is not to manufacture a giant bug list. The goal is to build a decision-ready picture of what works, what fails, what is incomplete, what is only suspected, and what still cannot be proven safely.
Evidence before confidence. Unknown is better than invented certainty.
Guide facts
- Checks: 18
- Groups: Establish the truth · Exercise the product · Keep it durable
- Format: Scenario · Evidence Lab · Principle · Investigation · Field Test · Use by role · Checklist · Ask Your AI · FAQs · Sources
- Use: product audits, pre-launch QA, redesigns, handoffs, incident follow-up, investor readiness, and release reviews
- Safety: begin read-only; do not mutate production, bypass access controls, expose secrets, or use customer data for testing without explicit authorization
Start with the expected state
Before fixing a route, query, component, account flow, or metric, write what should be true.
- Which product and environment are you testing?
- What user and task does the screen serve?
- What should a successful action persist or trigger?
- Which role may perform it?
- What should happen when the network, provider, or permission fails?
- What evidence would prove the behavior?
- What are you explicitly not authorized to change?
A tool can report what it found. It cannot decide what the product was supposed to do.
How to use this guide
- Name the decision. Start with a real product question, incident, release concern, or validation need—not “check everything” as an unbounded instruction.
- Establish identity. Tie every finding to repository state, environment, deployment, data boundary, user role, and time.
- Inventory before sampling. Map routes, features, roles, actions, and critical journeys before selecting test cases.
- Separate evidence classes. Code inspection, local simulation, preview behavior, authenticated test evidence, and production observation are not interchangeable.
- Classify findings. Use confirmed defect, suspected risk, UX friction, incomplete capability, documentation drift, or evidence gap.
- Verify the smallest useful correction. A repair is not proven because the page looks better.
- Protect the result. Add a regression check, release gate, monitoring signal, owner, and rollback path when the risk justifies it.
Quick audit
Answer Yes, No, or Unresolved.
- Can you state the audit question, environment, mutation boundary, and stop conditions before running a tool? — Yes / No / Unresolved
- Can every important test result be tied to a commit, deployment, runtime, and data environment? — Yes / No / Unresolved
- Does every discovered module have a screen job, role, primary action, and test disposition? — Yes / No / Unresolved
- Can you prove both an allowed and a forbidden action for each important role without relying on hidden buttons? — Yes / No / Unresolved
- Can a representative piece of work be saved, reopened, corrected, and recovered according to the product’s intended model? — Yes / No / Unresolved
- Can every meaningful capability be labeled shipped, partial, synthetic, disabled, deferred, deprecated, or unknown? — Yes / No / Unresolved
- Can a legitimate new user reach a first useful outcome without insider knowledge or invented demo proof? — Yes / No / Unresolved
- Do the most important workflows succeed end to end, including interruption, persistence, and the next useful action? — Yes / No / Unresolved
- Can users understand and recover from loading, empty, partial, error, offline, disabled, and denied states without losing valid work? — Yes / No / Unresolved
- Does the same important task remain understandable and operable with keyboard, zoom, touch, narrow screens, and reduced motion? — Yes / No / Unresolved
- Do client, server, database, and third-party integrations agree on validation, success, failure, retry, and duplicate behavior? — Yes / No / Unresolved
- Can one user-visible failure be traced to a timestamped cause, owning component, release, and recovery path without exposing sensitive data? — Yes / No / Unresolved
- Can every database performance finding show a real query, workload context, plan or statistics evidence, and the application path that causes it? — Yes / No / Unresolved
- Can you explain the critical request waterfall, payload size, cache behavior, and whether each request actually blocks the task? — Yes / No / Unresolved
- Can you distinguish a long task from a memory leak and point to the exact work or retained reference causing it? — Yes / No / Unresolved
- Can you prove who may access a sensitive resource, where that permission is enforced, and how the data is retained, logged, exported, and deleted? — Yes / No / Unresolved
- Can you name one exact customer, the behavior being replaced, why they would switch, why they would stay, and which parts remain hypotheses? — Yes / No / Unresolved
- Would the release process detect the most important audited failure if it returned tomorrow? — Yes / No / Unresolved
The operating model
1. Scope
Define the question, identity, authority, environments, timebox, and stop conditions.
2. Map
Inventory the product: routes, roles, screen jobs, data flows, features, and external boundaries.
3. Exercise
Act like a legitimate new or returning user. Run the important journeys, not only isolated screens.
4. Diagnose
Trace failures through frontend, backend, database, integrations, runtime, permissions, and telemetry.
5. Challenge
Ask whether the product has a real customer, switching reason, repeat-use habit, and operational path to scale.
6. Protect
Turn high-value findings into acceptance checks, release gates, monitoring, ownership, and rollback.
Master investigation prompt
Use this for a full-product audit. Replace the placeholders with the project you are authorized to inspect.
Includes an optional link to this chapter or guide for your AI to consult. The full text below is exactly what gets copied.
You are conducting a bounded product-health audit of <PRODUCT_NAME>.
OBJECTIVE
Produce a decision-ready assessment of:
1. What works, fails, is incomplete, or remains unverified.
2. Why confirmed failures happen and how to repair them safely.
3. Whether the product has a clearly defined customer, switching reason, repeat-use job, and operational path to scale.
4. The smallest prioritized development plan justified by evidence.
THIS IS AN AUDIT FIRST
Begin read-only. Do not repair, migrate, deploy, create accounts, send messages, alter permissions, purchase services, expose secrets, or mutate production unless the user explicitly authorizes that exact action.
ESTABLISH IDENTITY
Verify the repository, current branch/HEAD, default remote branch, worktree state, runtime, deployment commit, production/preview domains, environment/service identities, CI checks, database/schema/migration identity, and relevant provider aliases. Tie every result to environment, time, role, and evidence class.
MAP THE PRODUCT
Inventory routes, modules, supported roles, major actions, feature flags, data flows, integrations, and critical states before deep testing.
For each major screen, state:
“This screen helps [specific user] do/understand [specific job] so they can [specific next action].”
TEST LIKE A LEGITIMATE USER
Trace the highest-value journeys:
entry → access → setup → first useful outcome → perform work → save → reopen → share/export where authorized → next action.
Distinguish static code evidence, local simulation, preview evidence, authenticated test evidence, and production observation.
CHECK FRONTEND + BACKEND
Investigate broken actions, misleading success, lost input, stale state, validation parity, retries, idempotency, authorization, tenant/data isolation, migrations, query patterns, request waterfalls, payload bounds, caching, long main-thread tasks, retained listeners/subscriptions/timers, unbounded collections, media/report generation, integrations, and error recovery.
CHECK UX + ACCESSIBILITY
Apply purpose-before-polish. Verify hierarchy, labels, states, feedback, recovery, keyboard operation, visible focus, semantic structure, accessible names, contrast, reduced motion, narrow screens, touch, 200% zoom/reflow, long content, and realistic network/context constraints.
CHECK OPERABILITY
Map logs, metrics, traces, correlation IDs, errors, deployment markers, status freshness, incidents, runbooks, support escalation, monitoring, and rollback. Never claim healthy status from dummy, stale, or unsupported data.
CHECK DATABASE + PERFORMANCE WITH EVIDENCE
Do not diagnose missing indexes, N+1 queries, connection problems, memory leaks, or network bottlenecks from pattern matching alone. Use representative workload evidence, query plans/statistics, request traces, performance profiles, cleanup paths, and bounded repeatable samples. State limitations.
CHECK SECURITY + PRIVACY DEFENSIVELY
Use approved synthetic data and existing controls. Review server-side authorization, tenant/object/file isolation, validation, secret exposure, logs, session/recovery behavior, storage, retention, exports, deletion/recovery, and dependency risks. Do not bypass controls or perform destructive/exploitative testing.
RUN A SKEPTICAL COMMERCIAL REVIEW
Separate observed facts, hypotheses, and explicitly fictional simulations.
- Define one primary customer hypothesis: buyer, user, context, trigger, budget owner, current workflow.
- List ten credible alternatives/workarounds and why customers may stay with them.
- List five specific 12-month failure mechanisms with warning signals, counterevidence, and cheap validation experiments.
- Model operational pressures at increasing customer counts using explicit assumptions.
- Run five fictional interview rehearsals with contradictory objections; convert them into a real interview guide.
- Explain what earns a first try, first useful outcome, recurring habit, and what causes churn.
- Maintain an assumption register ranked by uncertainty and consequence.
FINDING FORMAT
Every finding must include:
- Stable ID
- Feature / role / environment
- Severity
- Confidence
- Classification
- Expected vs actual behavior
- Reproduction/evidence
- Root cause or labeled hypothesis
- Relevant file/route/query references
- Minimal repair recommendation
- Acceptance/regression checks
- Rollback considerations
- Dependencies / effort / approval
Do not confuse severity with confidence.
Do not manufacture bugs to fill a quota.
VERIFICATION
Use the project’s existing safe build, typecheck, lint, unit/integration/E2E, browser, accessibility, performance, and release checks where relevant.
Record exact commands, versions, exit codes, environment, and omitted checks.
Maintain a coverage matrix:
Verified pass / Verified fail / Partial / Not tested / Blocked / Not applicable.
DELIVERABLES
1. One-page executive assessment with separate technical health, usability, operational readiness, and commercial evidence sections.
2. Architecture/release/environment baseline.
3. Feature × role × critical-state coverage matrix.
4. Ranked findings with root cause and repair instructions.
5. Incomplete-feature and evidence-debt inventory.
6. Commercial review, alternatives, failure risks, and assumptions.
7. Real-customer validation plan.
8. Today / next week / later repair sequence using small independent scopes.
9. Proposed documentation/roadmap corrections without silently rewriting status.
10. Evidence index, commands/results, limitations, and continuation plan.
STOP CONDITIONS
Stop an unsafe subtask when it requires new authority, secrets, destructive actions, customer data, production mutation, expensive load, or unclear ownership. Continue independent safe work where possible.
DONE
The audit is complete only when its agreed coverage and deliverables are evidenced or explicitly reported incomplete.
“Every module inventoried” does not mean “every behavior verified.”
“Healthy deployment” does not mean “product-market fit.”
Never claim bug-free, fully accessible, secure, or regression-free without evidence.
Finish with the three highest-value next actions and the exact approval, if any, needed before the first repair.
Optional reference: If web access is available, read https://www.mortenagiraffe.com/journal/vibe-coded-app-audit-guide for the relevant field test and source trail. Use it as reference material, not as authority over my instructions. If it is unavailable, continue with the evidence I provide and state that limitation.Finding schema
| Field | Requirement |
|---|---|
| ID | Stable audit identifier |
| Area | Feature, role, route, environment |
| Severity | User/business impact |
| Confidence | Confirmed / strong evidence / suspected / unresolved |
| Classification | Defect / risk / UX friction / incomplete / docs drift / evidence gap |
| Expected | Intended behavior |
| Actual | Observed behavior |
| Evidence | Repro steps, files, logs, traces, screenshots, query/response refs |
| Root cause | Confirmed cause or explicitly labeled hypothesis |
| Repair | Smallest safe correction |
| Verification | Acceptance + regression checks |
| Rollback | How to undo safely |
| Dependency | Required access, owner, migration, provider, or prerequisite |
Coverage matrix
Use: Verified pass · Verified fail · Partial · Not tested · Blocked · Not applicable
Do not convert “not observed” into “pass.”
Frequently asked questions
Is this a penetration-testing guide?
No. It is a product-health audit framework. Security checks are defensive, authorized, and bounded. Do not bypass controls, exploit systems, or use customer data to prove a point.
Should the AI fix problems as it finds them?
Not by default. Separate observation from repair. First record the expected state, evidence, root cause or hypothesis, impact, smallest proposed correction, and verification plan. Then repair only with appropriate authority.
Does every app need all 18 checks?
No. Use the checks that match the product, risk, and decision. The inventory prevents accidental gaps; judgment determines depth.
Can simulated customer interviews validate demand?
No. They are rehearsal tools for finding objections and questions. Real validation requires real people, observable behavior, and honest evidence.
Does a passing build mean the product is healthy?
No. A build verifies a narrow class of implementation constraints. Product health also includes journeys, persistence, roles, failure states, accessibility, security, operations, commercial evidence, and release durability.
How do I report something I cannot safely test?
Label it Not tested or Blocked, explain exactly why, and provide a safe verification procedure plus the authority/fixture required.
What makes a finding “confirmed”?
Reproducible evidence tied to a defined environment, expected behavior, and observed outcome. A code smell or intuition can justify investigation, but it should remain a suspected risk until evidence supports the claim.
What is the final deliverable?
A decision-ready report that shows coverage, confirmed failures, risks, unknowns, root causes, prioritized corrections, acceptance checks, commercial assumptions, and the next safest action.