Short answer: AI coding tools are useful, and they are also wrong often enough that a person who knows the codebase must read every change before it ships. Our stated policy is simple: we use AI tools, and every change is reviewed by a senior engineer.
This post explains why we hold that position, what a reviewer should look for, and when you should not rely on it.
Is AI-generated code safe to put in production?
Not on its own. Treat it as a draft from a fast, confident junior who has never seen your production environment.
The public evidence points the same way. Veracode's 2025 GenAI Code Security Report found that AI-generated code introduced risky security flaws in 45% of its tests, across more than 100 language models and four languages (Java, JavaScript, Python and C#). It also reported that larger, newer models showed no meaningful improvement in security outcomes. An earlier academic study of GitHub Copilot, "Asleep at the Keyboard?", prompted the tool with 89 security-relevant scenarios and found that about 40% of the 1,689 generated programs contained vulnerabilities. That study dates from 2021 and tools have changed since, so read it as the origin of the concern, not as a current rate.
None of these studies describes your project. They describe a tendency, and the tendency is enough to justify a review step.
Do developers trust AI code themselves?
Mostly no, even while they use it. In the Stack Overflow Developer Survey 2025, 84% of respondents said they use or plan to use AI tools. In the same survey, 46% said they distrust the accuracy of AI output and 33% said they trust it. The top frustration, named by 66%, was AI solutions that are "almost right, but not quite". And 45.2% said debugging AI-generated code takes more time.
"Almost right" is the dangerous category. Code that fails loudly gets fixed. Code that works on the happy path and fails on the third unusual input gets shipped.
Why a senior engineer, and why every line?
Because the mistakes that matter are the ones that need context to see. A reviewer has to know what the system is supposed to do, what data it touches, and what the team has already decided not to do. Experience is what makes "this looks plausible" turn into "this is wrong here".
Reviewing every change, not a sample, is a stance. A sampled review tells you about the average change. A breach or a billing bug comes from the one change nobody read.
This does not mean the AI tool is banned or that the reviewer retypes everything. It means the person who approves a change can explain it, and is accountable for it.
What should a reviewer check in AI-generated code?
Use this as a checklist. It is the stance we would hold ourselves to, written down so you can ask any studio the same questions.
Security
- Is all user input validated and escaped before it reaches a database query, a shell command, a file path or a page?
- Are authentication and permission checks present on every route that needs them, not just the obvious ones?
- Are passwords, tokens and sessions handled with the framework's own tools rather than hand-rolled code?
- Does the code log anything it should not, such as personal data or tokens?
Dependencies
- Is every new package real, maintained and actually needed? Research on package hallucination found invented package names in at least 5.2% of samples for commercial models and 21.7% for open-source models, across 576,000 code samples. An invented name that someone later registers with malicious code is a supply-chain risk.
- Could the same job be done with what the project already uses?
- Is the version pinned, and does the package's own dependency list look reasonable?
Tests
- Does the change come with tests that would fail if the behaviour were wrong?
- Were the tests written to check the requirement, or merely to pass against the code as written?
- Did a person run the code and see it work, rather than trusting the summary?
Licences
- Does any new dependency have a licence compatible with how the client will use and distribute the software?
- Does a large block of code look copied from somewhere identifiable? If so, find where from and check its licence.
Secrets
- Are there any keys, passwords or connection strings in the code, in examples, in comments or in test data?
- Do configuration values come from environment settings and not from the repository?
Edge cases
- What happens with empty input, very long input, duplicates, missing data, a failed network call, a time zone change, a second click?
- What does the change do to existing data? Can it be undone?
- Does it behave properly in both languages on a bilingual site?
If a change cannot be explained in a sentence by the person approving it, it is not ready.
What do you gain from a studio that uses AI tools?
Possibly speed on routine work, such as boilerplate, tests for existing behaviour and first drafts of repetitive code. Whether that becomes a saving for you depends on the project and on how much review effort the change needs. We would not promise a percentage, and you should be cautious of anyone who does without showing how they measured it.
Questions worth asking any studio:
- Do you use AI tools on client work, and for what?
- Who reviews the output, and is every change reviewed or only some?
- How do you keep client code, data and keys out of tools you do not control?
- Who is accountable when a change causes a fault?
- Do I get the repository and the history of reviewed changes?
On the third question, check the terms of the specific tool in use. Data handling differs between products and plans and changes over time, so read the vendor's current policy rather than relying on a general answer.
When this is not for you
This approach adds review time, and review time is real work. If you want the cheapest possible prototype to test an idea for a weekend, unreviewed AI output may be a reasonable choice, as long as you treat it as disposable and never connect it to real customer data or payments. We would not recommend taking that prototype live as it stands.
Likewise, if your software handles regulated data, payments or health information, a review policy is a starting point, not a complete compliance answer. You will need security testing and legal advice on top.
Next step
If you are weighing a website, app or online store and want to know how the work would be reviewed before it ships, ask. You will talk to Kevin directly, and he will answer the five questions above for your project.



