How to review AI-generated code before you merge it
Fluent code is not correct code. An AI patch is a suggestion that still has to survive the same review you would give a rushed teammate.
Deni AI team
Review the patch, not the confidence
AI-generated code often arrives with comments, a migration plan, and a calm explanation of why the change is safe. That packaging is not evidence. The evidence is a diff you can run.
We treat generated patches the way we treat a pull request from someone who has never seen the repo: assume they guessed the boundaries, then prove otherwise.
Did it touch the right files?
Ask the model to name files before it writes a patch. If the answer wanders into unrelated modules, treat the whole draft as a sketch. Scope errors are cheaper to catch than logic errors.
Can you run it immediately?
If you cannot paste the change and run tests, types, or the app, you do not have a patch. You have a description of a patch. Do not review prose as if it were a diff.
What did it invent?
Scan for new helpers, flags, env vars, and endpoints that were not in the excerpt you pasted. Invented APIs are the default failure mode of coding models, not a rare bug.
Is it faster to rewrite?
If the draft fights the existing style, ignores tests, or requires a paragraph of cleanup per file, throw it away. Keeping a bad AI patch out of loyalty wastes more time than starting from the failing test.
A review order that stays cheap
First, confirm the intended behavior in one sentence. If you and the model do not share that sentence, stop. Prompting for more code will only decorate the misunderstanding.
Second, look at file list and public API. A good patch is boring: it changes the smallest surface that can carry the behavior. A bad patch refactors neighbors to make the new idea fit.
Third, run the smallest check that can fail: a unit test, a typecheck, or the one screen the change affects. Reading without running is how invented helpers survive.
Security is not a later pass
Watch for new network calls, loosened auth checks, logged secrets, and copy-pasted snippets that pull in a dependency you did not ask for. Models optimize for “it works in the story,” not for your threat model.
If the task touches auth, billing, or user data, the human review is the product. The model can propose a patch. It cannot accept the risk.
Merge checklist
- The behavior is stated in one sentence you agree with.
- The file list is small and named before the patch.
- No new API, flag, or env var appeared without a source in the repo.
- Tests or a manual path were run, not only described.
- You would still understand the change if the chat disappeared tomorrow.
How this fits a multi-model workspace
In Deni AI we start with a coding-capable model, then switch only if the first answer cannot explain its constraints. The workspace is for that switch. It is not a substitute for the typechecker.
If you want the broader verification method for facts and citations, not only code, use the verify-AI-answers guide. The habits are the same: name the failure, then check it outside the chat.
Common questions
Should I ask the model to write the tests too?
You can. Then run them. Tests generated with the same guess can share the same blind spot. Prefer a test you understand over a green suite you cannot explain.
Is a coding model enough, or do I need a second model?
Use a coding-capable model for the first patch. Bring a second model only when the task is ambiguous or the first answer cannot name its constraints. Running two models does not replace running the code.
When is AI code not worth reviewing?
When you cannot describe the expected behavior, or when the change is one line you already know. Review cost should not exceed the cost of writing it yourself.