Skip to content
Insights

NQ Solution

How to Review AI-Written Code (2026)

A practical order for reviewing AI-written code: scope, intent, tests and behaviour, the parts that always need a person, and what to ask your developer.

To review AI-written code, check four things in order: that the change stays within what was asked, that it matches the intent of the system, that tests prove the behaviour rather than just pass, and that a person has read every part touching permissions, data or money. The question became more pressing in September 2026. Anthropic released Claude Opus 5.5 on 22 September and Sonnet 5.5 on 28 September, and Linear wrote on 21 September that "agents now write the majority of our tests". When an agent produces most of the code, review is where quality is decided.

We build software with an AI coding agent every day, so this guide reflects how we work, plus the questions to ask if someone else is building for you.

Why is reviewing AI-written code different from reviewing human code?

A human developer's mistakes usually follow their understanding: if they misread the requirement, the whole change reflects that. An AI agent's mistakes are less predictable. The code is often tidy, well named and confidently commented, which makes it look reviewed when it is not.

Three patterns come up again and again:

  • Scope drift. You ask for one change and get five. A satirical site that reached the top of Hacker News on 9 September 2026 was built around a single prompt: "Claude, change the 'Add to Cart' button to blue." The joke landed because people recognised it: small requests can return large, unrelated edits.
  • Plausible but wrong. The code compiles, the tests pass, and the behaviour is still not what the business needs, because the agent filled a gap in the instructions with a reasonable guess.
  • Known vulnerability patterns. Veracode's 2025 GenAI Code Security Report found that 45% of the AI-generated code samples it tested introduced an OWASP Top 10 vulnerability. Veracode's March 2026 follow-up, as summarised by the Cloud Security Alliance, found the security pass rate still flat at about 55%, and missing authorisation checks and unsafe input handling remain typical failures.

A working definition: AI code review is checking machine-written changes for scope, intent, data access and behaviour before they ship, not only for syntax.

What should you check first: scope, intent, or tests?

Scope first. It is the quickest check and it changes how you read everything else.

  1. Scope. Compare the diff with the request. List every file changed. Anything outside the request needs a reason, or it comes out. Dependency changes, configuration edits and deleted tests deserve particular attention.
  2. Intent. Does the change fit how the system is meant to work? This is where written project instructions pay off. Many teams keep an AGENTS.md or CLAUDE.md file in the repository that tells agents the architecture, conventions and rules ("never write to the orders table directly"). Claude Code and other agents read these files. If the change breaks a written rule, the review is short.
  3. Tests. Read the tests before trusting them. An agent asked to make tests pass can weaken an assertion, mock away the part that fails, or skip a test. Check that each new test would fail if the feature were broken.
  4. Behaviour. Run the thing. Open the page, submit the form, log in as the wrong user. A diff cannot show you a layout that collapses on a phone.

Linear's post shows the same idea at team scale. With agents writing most of their tests and the suites "almost quadrupling" since the start of the year, they updated their agent skills "so generated tests follow the same constraints by default". The rules for agents are written down, not left to each review.

Which parts of an AI-built app always need a human reviewer?

Some areas are cheap to get wrong and some are not. We let automated review handle style and simple bugs everywhere, but a person reads every line in these areas:

AreaWhat can go wrongWhat the reviewer checks
Authentication and permissionsA route or API endpoint is reachable without the right roleEvery new endpoint checks who is asking, on the server
Personal dataData is logged, exposed in an API response or stored without needWhat is collected, where it goes, who can read it
Payments and moneyAmounts are calculated on the client, or webhooks are trusted without verificationServer-side totals, signature checks, idempotency
Database changesA migration drops or rewrites dataThe migration on a copy of real data, and a rollback plan
Secrets and configurationKeys end up in client code or the repositoryEnvironment variables, build output, public files
DependenciesA new package is added for a small taskWhether it is needed, maintained and safe
Deletion and bulk actionsOne click affects many recordsConfirmation, limits and audit records

The rest of the app (layout, copy, simple components) can move faster with lighter review, as long as someone looks at the result on screen.

Can AI review AI code? (and where it falls short)

Yes, and it is worth doing. Automated review tools, including Claude Code's /code-review command, read a change in the context of the whole repository and flag bugs, missed edge cases and breaks from project conventions. Anthropic's Opus 5.5 announcement quotes Deloitte Consulting: even at its lowest effort setting, the new model "caught 72% of known bugs in our code reviews", against 56% for Opus 5 at high effort. That is a useful first pass, and it is also a reminder that more than a quarter were missed.

Where AI review falls short:

  • It shares the author's blind spots. If the requirement was ambiguous, the reviewer model may accept the same wrong guess.
  • It does not know your business. It cannot tell that a discount must never stack with another, unless that rule is written somewhere it can read.
  • It cannot see the running system. Production settings, real data and how users actually behave are outside the diff.

So we use AI review to raise the floor, then a person decides. A finding from the tool is a question to answer, not an instruction to follow.

Who owns code written with AI tools?

There are two separate questions: copyright law and your contract.

On copyright, the US Copyright Office's January 2025 report on copyrightability concluded that protection depends on human authorship. Material generated by AI without sufficient human control over its expressive elements is not protected, while human selection, arrangement and modification can be. Most commercial code involves a lot of human direction and editing, but the law is still developing and differs by country. Take legal advice if ownership is central to your business.

The contract question matters more day to day. Your development agreement should say who owns the delivered source code and other outputs, and on what conditions they are handed over, whether or not AI tools were used. We set this out in the contract for each project.

What should you ask a dev studio that builds with AI?

If you commission software, you do not need to read the code yourself. Ask these questions instead, and look for specific answers:

  1. Which AI tools do you use, and for what? "We use them for everything" and "we never use them" are both answers worth probing.
  2. Who reviews each change before it ships? You want a named person, not "the tool checks it".
  3. What gets human review every time? Listen for permissions, personal data, payments and database changes.
  4. How do you test it on the real screen? Ask what they check in a browser before handover.
  5. What do we receive at the end? Source code, documentation and the conditions for handover should be in the contract.

For more on choosing a team, see our guide to outsourcing software development to South Korea.

How we handle this on client projects

We develop and review code every day with Claude Code, an AI coding agent, and a person reviews every change before it ships. In practice that means:

  • Claude Code reviews changes as well as writing them, and its findings are a first pass, not the decision.
  • A person reads and approves each change before it is merged and released.
  • We build with security checks as part of the work: permissions, input validation, dependency patching and checks for what the app exposes publicly.

We do not claim AI makes us a fixed amount faster. The agent writes and suggests; we decide. If you are running AI-generated workflows inside your business rather than in code, the same principle applies; see AI workflow automation with human review. For apps built quickly with AI by someone else, our vibe-coded app security checklist lists the gaps we see most often.

Frequently asked questions

How do small teams review AI-generated code without slowing down?

Make changes small, write the project rules down where the agent reads them, and let automated review catch style and simple bugs. Spend human attention on the areas that are expensive to get wrong: permissions, data, money and database changes.

Is AI-generated code copyrighted?

In the US, the Copyright Office's January 2025 report says protection requires human authorship. Purely AI-generated material is not protected, while human-directed and edited work can be. Ownership between you and your developer should be set out in the contract either way.

What is AGENTS.md and why does it matter?

AGENTS.md is a plain text file in a repository that gives AI coding agents instructions: how the project is structured, which commands to run and which rules to follow. Agents such as Claude Code read it, so written rules are applied to every change instead of being repeated in each prompt.

Do AI code review tools replace human code review?

No. They are good at finding many bugs quickly, but they share blind spots with the model that wrote the code and do not know your business rules. Use them as a first pass, then have a person approve.

How do I know if my developer used AI?

Ask. Most professional developers now use AI tools for some of their work, so the useful question is not whether they use them but who reviews the result and how it is tested.

Sources

All checked October 2026.

Building something with AI in the loop?

We build with Claude Code every day and review every change before it ships. Tell us what you are building through the project request form, or email dwkim@nqsolution.kr. Our software development in Korea page explains how we work with English-speaking clients.

Planning a similar project?

Send us your goals, scope and timeline. The two of us who would build it reply directly.

Prefer email? Write to dwkim@nqsolution.kr — the founder replies directly.