An AI coding tool can produce an app that looks finished. Someone still has to establish what the app does, what data it touches and whether it is safe to use. The first check is not whether the screen looks polished; it is whether the code and the system around it behave as intended under realistic conditions.
In June 2026, GitHub said its security validation for third-party coding agents had become generally available. It checks for potential vulnerabilities, risky dependencies and exposed secrets in supported repository workflows. Those checks are useful, but their presence is not a guarantee that an app is secure or suitable for a particular business.
Know what the generated app can reach
List the services, files, databases and accounts the app uses. If a simple booking page has access to the whole customer database, the permission is wider than the task. If the app stores email addresses or payment details, understand where they go and how they are protected. Ask the developer to show the actual configuration and the path a user's data takes.
Do not put passwords or API keys into source files to make a demo work. Check the repository for accidentally committed secrets and rotate anything already exposed. GitHub describes secret scanning as one of its controls for agent-created changes; other hosting setups need equivalent checks.
Read the dependencies and the parts that accept input
AI may suggest packages that are outdated, unnecessary or not the package the developer intended. Review the dependency list, lockfile and update policy. An app that accepts form submissions, uploaded files or external data needs validation before that input reaches a database, page or privileged action.
OWASP's 2025 guidance on improper output handling explains how untrusted model output passed into other systems can lead to security problems. The same discipline applies to generated code: treat apparently plausible output as material to inspect and test, not as a trusted implementation merely because it runs.
Run tests that reflect the real task
A screenshot proves that one state rendered. It does not show what happens with an empty field, a long name, duplicate submissions, an interrupted payment, a user without permission or a lost network connection. For a booking form, try a valid booking, an invalid date, two people requesting the final slot, and a retry after a timeout. Confirm what the user sees and what the system recorded.
This is a suggested test plan, not a claim that every generated app has these faults. Choose scenarios from the app's actual purpose. Automated unit tests can catch repeated mistakes; an independent human review is still valuable for security, accessibility and the meaning of the user flow.
Decide who signs off before release
Give one person responsibility for approving the change and another for operating the service after launch if the team allows. Record the code revision, the tests run and any known limitations. If the app handles sensitive information or a consequential decision, use a developer or security reviewer with relevant expertise. A small prototype can be explored privately; production use asks for a higher standard.
Make rollback and incident handling part of the release plan. If a bug appears after launch, can you remove the feature, restore data and tell affected users what happened? “The AI wrote it” does not answer those questions.
A five-question handover
Ask for a small demonstration using test data. A reviewer should be able to follow a user request from the interface through the server and database, see what is stored, and verify what appears after a failure. They should also be able to show where logs go and who can access them. This is particularly helpful for a nontechnical owner because it turns an abstract claim of “secure code” into observable behaviour. If the developer cannot explain a surprising package or permission the AI added, remove or justify it before release. Keep the review tied to the actual version you intend to deploy; a later AI edit can change the answer.
Before accepting an AI-built app, ask: What does it access? Where are the secrets? Which dependencies are required? Which failure cases were tested? Who can fix it after launch? If the answers are unclear, the app is still a prototype. The useful speed of AI coding comes from making a working draft faster; safe deployment still requires evidence about the finished system.
Sources & further reading
- GitHub: Security validation for third-party coding agents — checked 2026-10-08
- OWASP: Improper output handling — checked 2026-10-08
