Your AI coding tool says the tests passed. GitHub shows a green check. The app opens on your laptop.
Can you invite people to use it?
Maybe. But those three facts do not tell you whether saving works, whether another account can read private records, or whether your tests ran against the version you are about to publish.
I want beginners to ask a more useful question than "Is everything green?"
What actually ran, and what does its result prove?
On October 1, GitHub changed its Code Quality coverage-upload action. A push to a non-default branch without an open pull request now skips the upload instead of failing because the pull request number is missing. A notice and step summary explain the skip. Supported pull request events and default-branch uploads keep their existing behavior.
There is a timing detail: a workflow triggered only by pushes will not upload merely because you open a pull request. It needs another push, or an appropriate pull request trigger. This particular feature is for GitHub Team and Enterprise Cloud, not Enterprise Server.
That is a sensible fix for a misleading failure. My broader lesson is not that green checks are bad. It is that a job completing successfully and a particular piece of evidence being collected are different facts.
Start with four honest outcomes
I would give every important check one of four labels:
- Passed: the check ran, and its stated expectation was met.
- Failed: the check ran, and its expectation was not met.
- Skipped: the check intentionally did not run, with a recorded reason.
- Unknown: there is no reliable result, or the result cannot be matched to this build.
A skipped coverage upload does not mean the tests failed. It also does not mean coverage was uploaded. Both statements can be true without any contradiction.
The dangerous move is translating skipped or unknown into passed because the overall page looks reassuring.
Imagine a hypothetical appointment app. Its automated build succeeds, but the check that needs a test database is skipped because that database was unavailable. The app may still package correctly. You have not established that it can save an appointment.
If you are just getting started, the $1 AI App Builder Starter Prompts guide your AI from an idea toward a working first build. Keep this four-outcome rule beside those prompts: an assistant should explain what it could verify, not hide missing checks behind a confident summary.
Give each check a narrow job
"Test the app" is too broad to judge. I would break it into claims a beginner can understand.
A build check asks whether the source can be turned into a runnable product. A small logic test might ask whether cancelling an appointment changes its status. A browser test might ask whether a user can book a slot and see it after refreshing.
A permission test asks a different question: can a second user read or change an appointment they do not own?
One successful check does not substitute for all the others. An app can compile while saving nothing. A save can work while ownership protection is broken. A correct permission rule can exist while the screen still sends the wrong request.
Before asking AI to add tests, write the behavior you want to protect in plain English:
"A signed-in customer can book an available slot, reload the page, and see the booking. A different customer cannot view or cancel that booking."
Now your assistant has a concrete promise to test rather than a vague request to produce impressive output.
Read the denominator before the percentage
Code coverage describes which portions of code were exercised during a test run. It is useful evidence, not a verdict on whether the product is good.
For example, Vitest's documentation explains that its default report shows files imported during the test run. To include uncovered files, you need an appropriate inclusion pattern.
That makes the denominator important. A high percentage across a small selected portion of your project is not the same claim as a high percentage across the application you intended to assess.
Ask your AI which source files the report includes and excludes. Ask whether the important save, permission, and recovery paths are represented. Do not demand an arbitrary percentage before understanding what it measures.
Even fully exercised code can contain weak assertions. A test that clicks Save and checks only that the button exists proves much less than a test that reloads and checks the saved appointment.
Require evidence from the exact build
Yesterday's successful test run is a useful history entry. It is not automatically evidence for today's edited version.
I would ask AI to keep a small verification note with each candidate release:
- Version: the commit identifier, or another exact build identifier.
- Environment: local development, test deployment, or production-like staging.
- Checks: names, outcomes, and reasons for any skips.
- Evidence: paths to reports, screenshots, or relevant logs.
- Gaps: behavior that remains untested or could not be verified.
- Decision: what can proceed and what must wait.
Keep credentials, private customer records, and access tokens out of that note. Test with synthetic records whenever practical.
This is not paperwork for its own sake. It prevents a screenshot from an old preview or a report from a different branch from quietly becoming approval for the current app.
Test the journey, not just the machinery
Testing Library's guiding principles favor checks that resemble how people use the interface rather than relying on a component's private internals.
For our hypothetical appointment app, I would start with one complete journey: sign in, choose a slot, book it, reload, inspect it, and cancel it. Then I would test the same record with another account and an invalid request.
Use a disposable test environment for destructive actions. Do not run a cancellation test against someone's real booking.
This approach also helps you review AI work. You can understand "the appointment survives a refresh" without understanding every line of its storage implementation.
The $9 AI App Builder From Zero e-book connects these steps to the wider path from idea to publication: scope, design, architecture, building, testing, and launch. All 40 Starter Prompts are included as a free bonus inside its PDF and EPUB; you are paying for the e-book, not an extra prompt pack.
The prompt I would use before a beta
Inspect the current project and identify its exact build or commit. List the verification checks that actually ran for this version. Classify each as passed, failed, skipped, or unknown. Explain what each result proves in plain English. Show the report or evidence path, and identify excluded files, mocked dependencies, missing environments, and untested user journeys. Do not convert a skip into a pass. Propose the smallest safe test that would resolve each important unknown. Do not change production data, spend money, or deploy without asking me.
Read the answer for missing evidence, not just positive language.
If a required check is skipped, restore its prerequisite or run an appropriate equivalent and record the result. If you cannot do either, keep that limitation visible and defer the affected release decision. Do not remove the check just to make the dashboard green.
Review Radar is coming soon. Its planned package brings review-backed research, tailored screen designs, and an AI-readable project folder together with build prompts and acceptance criteria. View the preview and join the email waitlist. Checkout is not open; the package does not promise a finished app, users, or revenue.
A green dashboard should be the beginning of your evidence review, not the end of your thinking.
You can also find me here:
Medium: https://medium.com/@marcusykim
DEV.to: https://dev.to/marcusykim
Website: https://marcusykim.com/
X: https://x.com/marcusykim
LinkedIn: https://www.linkedin.com/in/marcusykim/
Upwork: https://www.upwork.com/freelancers/marcusykim
Top comments (0)