SecurityBlog Post

How to Test a Vibe-Coded App Before Launch: A Practical Vibe Testing Guide

Clicking through your app and seeing it work is not testing. Here is what vibe testing actually means, the specific tests that catch the bugs AI tools leave behind, and how to stop AI-written tests from quietly confirming those bugs.

By Abhay Mittal

9 min read
How to Test a Vibe-Coded App Before Launch: A Practical Vibe Testing Guide
Is your AI-built app exposed? Get a professional vibe coding audit and ship to production with confidence.

Most vibe-coded apps are tested the same way: the founder clicks through the main flow, everything works, and the app goes live. That checks the happy path. Real users, and attackers, do not stay on it.

This guide covers how to test an AI-built app properly without becoming a QA engineer. It focuses on the tests that find the problems AI coding tools most often introduce: broken access control, unhandled edge cases, and payment flows that can be skipped.


What is vibe testing?

"Vibe testing" is used in two ways. Some people use it for testing by feel: clicking around until the app seems fine. Others use it for testing a vibe-coded app specifically, often with AI writing the tests too.

The first meaning is the problem. The second can work well if you know where AI-written tests go wrong. In this guide, vibe testing means a lightweight, structured way to test an AI-built app that a non-specialist can run, with a handful of high-value checks rather than a full test suite.


Why AI-built apps need a different testing focus

AI coding tools are good at making features work. The code usually compiles, renders, and handles the inputs you described. The gaps tend to be:

  • What happens for the wrong user. Routes check that someone is logged in, not that the data belongs to them.
  • What happens with unexpected input. Empty strings, negative numbers, huge files, repeated clicks.
  • What happens when a step is skipped. Paid features unlocked by visiting the success URL directly, or onboarding steps that can be bypassed.
  • What happens at the edges of time. Expired trials, cancelled subscriptions, reused password reset links.

So the testing priority is not "does it work?" It is "does it refuse to work when it should?"


Test 1: The two-account access test

This single test finds the most common serious bug in vibe-coded apps. It takes about fifteen minutes.

  1. Create two accounts, A and B, in separate browsers (or one normal window and one private window).
  2. As user A, create some private data: a project, an invoice, a document, a message.
  3. Open the browser dev tools Network tab and find the request that loads that data. Note the URL and any IDs.
  4. As user B, try to load the same URL or replay the request with B's session. In Chrome you can right-click a request and copy it as fetch or cURL.
  5. Try editing and deleting A's data as B the same way.

User B should get a 403 or 404 every time. If B can see or change A's data, you have an insecure direct object reference (IDOR). Our deep dive on the IDOR bug in AI-generated API routes shows how to fix it in code.

If your app uses Supabase directly from the browser, also run this against the database. The Supabase row-level security guide explains how to query as a specific user.

If you have teams or organisations, repeat the test across teams: an admin of team 1 should not be able to touch anything in team 2.


Test 2: Roles and permissions

If your app has roles (member, admin, owner, staff), list what each one should and should not be able to do. Then, for each "should not":

  • Log in as the lower role.
  • Find the request the higher role makes (for example, deleting a user or changing a plan).
  • Send it as the lower role.

Hidden buttons are not access control. The check has to happen on the server. The OWASP Authorization Cheat Sheet is a good reference for what correct enforcement looks like.


Test 3: Edge cases and bad input

Go through every form and API input and try:

  • Empty values and whitespace only.
  • Very long strings (paste a few thousand characters).
  • Negative numbers, zero, and decimals where whole numbers are expected.
  • Quantities and prices you edit in the request rather than the UI.
  • Duplicate submissions: click the submit button rapidly or replay the request.
  • Files of the wrong type or a very large size.
  • Special characters: quotes, angle brackets, emoji, and non-Latin text.

You are looking for crashes, data that should have been rejected, and anything that renders raw HTML. A checkout that accepts a quantity of -1 or a price from the browser is a revenue bug. See checkout price tampering.


Test 4: Authentication flows

Auth flows are generated quickly and tested rarely. Check:

  • Password reset links stop working after one use and after a time limit.
  • Email verification is actually required before access, not just sent.
  • Logging out invalidates the session. Copy a request, log out, replay it.
  • Rate limits exist on login and OTP endpoints. Try twenty wrong passwords quickly. See login and OTP rate limiting.
  • Changing email or password requires the current password or re-authentication.

Test 5: Payments in test mode

Never test payments with live keys. Stripe's testing documentation provides test card numbers for success, declines, and 3D Secure, and the Stripe CLI can forward webhooks to your local app.

Test these scenarios:

  1. Successful payment grants access only after the webhook arrives.
  2. Visiting the success URL directly without paying does not grant access.
  3. A declined card leaves the user on the free plan.
  4. Cancelling a subscription removes access at the right time.
  5. A failed renewal is handled, not ignored.
  6. A forged webhook (sent with curl and no valid signature) is rejected.

The last one matters more than it sounds. Our post on Stripe webhook signature bypass explains why unsigned webhooks let anyone grant themselves a paid plan.


Test 6: Check the AI-written tests

Asking Claude, Cursor, or another tool to "write tests for this feature" is a reasonable step. Read what it produces, because AI-written tests often have one of these problems:

They assert the current behaviour, including bugs. If the code lets any user fetch any invoice, a generated test may check that fetching an invoice returns 200, and pass.

They mock away the thing that matters. A test that mocks the database and the auth layer proves the function calls a mock correctly. It says nothing about access control.

They only cover the happy path. Ten tests that all log in as the owner and fetch the owner's data look thorough and test nothing about security.

Ask for the negative cases explicitly:

// The test you want: user B must not read user A's invoice.
test("other users cannot read an invoice", async () => {
  const invoice = await createInvoice({ ownerId: userA.id });
  const res = await request(app)
    .get(`/api/invoices/${invoice.id}`)
    .set("Authorization", `Bearer ${userBToken}`);
  expect([403, 404]).toContain(res.status);
});

A good prompt is: "Write tests that prove a user cannot read, edit, or delete another user's records, and that a non-admin cannot call admin endpoints. Do not mock authentication."


Test 7: Smoke tests with Playwright

Once the core flows work, a few end-to-end tests keep them working as AI tools keep changing code. Playwright runs a real browser and is easy to set up in a JavaScript project.

You do not need many. Five smoke tests cover most of the risk:

  1. Sign up and verify email.
  2. Log in and log out.
  3. Create, edit, and delete the main object in your app.
  4. Upgrade to a paid plan in Stripe test mode.
  5. Confirm a logged-out visitor is redirected away from the dashboard.

Run them before every deploy. When an AI change breaks login at 11pm, you want to find out before your users do.


Test 8: Basic load and cost checks

You do not need a full load test before launch, but you should know what happens under a little pressure:

  • Send a burst of requests to your most expensive endpoint, especially anything that calls an AI model. Is there a rate limit, or does your API bill grow without bound?
  • Check that list endpoints paginate. An endpoint that returns every row will slow down as you grow.
  • Confirm error pages do not show stack traces or environment details.

A one-hour vibe testing checklist

If you only have an hour before launch, do these in order:

  1. Two-account access test on your main data (15 minutes).
  2. Role escalation test on one admin action (10 minutes).
  3. Visit the payment success URL without paying, and send an unsigned webhook (10 minutes).
  4. Twenty failed logins in a row (5 minutes).
  5. Search the deployed JavaScript bundle for sk_, service_role, and API keys (5 minutes).
  6. Submit negative numbers and edited prices to checkout (10 minutes).
  7. Log out and replay a request (5 minutes).

Then go through the full vibe-coded app security checklist, or take the free vibe code security check for a scored summary.


When testing is not enough

Self-testing catches a lot. It is limited by what you think to test. A professional reviewer reads the code and tests the running app with your business rules in mind, which finds the paths you did not imagine.

If you are about to take real payments, handle sensitive data, or answer an enterprise security questionnaire, a code audit or penetration test is the next step. Our guide on how to review vibe-coded apps explains how to decide which you need.


FAQ

What is vibe testing?

Vibe testing usually means testing an app informally by clicking around until it seems to work, or testing a vibe-coded app with AI-generated tests. A useful version adds a few structured checks: two-account access tests, edge cases, auth flows, and payment tests.

Can AI write tests for my vibe-coded app?

Yes, and it saves time. Review the tests to make sure they check that unauthorised actions fail and that they do not mock away authentication or the database.

What is the most important test for an AI-built app?

The two-account access test. Logging in as one user and trying to read or change another user's data finds the most common serious vulnerability in AI-generated code.

Do I need automated tests before launch?

A handful of end-to-end smoke tests for sign-up, login, your core feature, and payments are worth having. A full test suite can come later.

How do I test vibe-coded apps for reliability?

Test edge cases and bad input, failed and cancelled payments, expired sessions, and bursts of requests. Then add smoke tests that run before every deploy so new AI changes do not break core flows.

Keep reviewing your app

Practical checks for the parts of an AI-built app that handle real users and money.

Need a second pair of eyes? Explore our code audit services, scope and pricing, and client case studies.

VibeAudits

Security Experts

Worried your vibe-coded app has issues like this?

We run professional code audits for SaaS apps and AI features built with Cursor, Claude, Copilot, Lovable and Replit. We find the security and reliability problems before your customers (or attackers) do, then hand you a fix-ready report.