1. Home
  2. AI governance
  3. Technical deep dive

Technical deep dive

Shipping AI-written code without shipping its bugs

How we let your developers keep the speed of AI coding assistants and agents while stopping more bugs before production: the checks every pull request must pass, review rules, safe releases, and clear rules for the AI tools themselves.

  • For CTOs & engineering leads
  • Read 8 min
  • Works with GitHub · GitLab · Azure DevOps
01 · The problem

Code is written faster than it is checked

AI coding assistants and agents let one developer produce far more code in a day. Review time, test suites and release processes didn’t speed up to match. Pull requests get bigger, reviewers skim, and the code looks right, which is exactly what makes its mistakes easy to miss.

AI also brings new kinds of mistakes:

  • Invented dependencies. A suggested package or API that doesn’t exist, or a look-alike name that an attacker can register.
  • Tests that prove nothing. Tests written to pass against the code, rather than to check what it should do.
  • Confident edge-case bugs. The happy path works, and the error handling, time zones and permissions don’t.
  • Data in prompts. Secrets, customer records or internal code pasted into tools that were never approved.

We build with AI every day at DevOpsPlant, and we govern it. The answer isn’t to slow developers down or ban the tools. It is to make every change prove itself automatically, and keep humans on the decisions that matter.

02 · Checks

The checks every pull request must pass

These run on every pull request, however the code was written. A developer gets the result in minutes, before a reviewer spends any time on it.

CheckWhat it catchesBlocks the merge when
Tests on changed linesNew code with no tests, or tests that don’t exercise itCoverage of the changed lines is below the agreed minimum
Security scanInjection, unsafe deserialisation and other common flawsA new finding at high severity or above
Dependency checkPackages that don’t exist, look-alike names, known vulnerabilities, unapproved licencesA new package isn’t on the approved list
Secret scanningKeys, tokens and passwords in the diffAny secret is found
Change sizePull requests too large to review properlyThe diff passes an agreed size, unless a lead approves the exception

Coverage of the changed lines matters more than overall coverage. A codebase can sit at 70% while the new, AI-written code has no tests at all.

.github/workflows/pr-checks.yml
      # Fail the pull request if the lines it changes are under 80% covered
      - name: Coverage of changed lines
        run: |
          pip install diff-cover
          diff-cover coverage.xml --compare-branch=origin/main --fail-under=80
03 · Review

Humans still approve

AI review tools are useful as a first pass, flagging obvious problems before a person looks. They don’t approve changes.

  • Every merge needs a human approval. Branch protection enforces it, including for changes an agent opens.
  • Sensitive code has named owners. Payments, authentication, infrastructure and the pipeline itself need a reviewer from the owning team.
  • Small pull requests by default. A reviewer can properly check a few hundred lines, not a few thousand.
.github/CODEOWNERS
# Changes here always need a reviewer from the owning team, however the code was written
/src/payments/   @acme/payments-leads
/src/auth/       @acme/security
/infra/          @acme/platform
/.github/        @acme/platform
04 · Release

Assume some bugs will still get through

No set of checks catches everything, so releases are built to limit the damage:

  • Feature flags. New features ship switched off, then turn on for a small group first.
  • Canary releases with automatic rollback. A new version takes a small share of traffic, and rolls back by itself if error rates rise.
  • Fast rollback. Going back is one step, practised, and faster than debugging in production.

The release side is covered in detail in From commit to production in minutes.

05 · Governance

Rules for the AI tools themselves

Governance is a short, written policy your developers can follow, enforced where it can be:

  • Approved tools. A list of the assistants and agents the team may use, on business accounts with training on your code turned off.
  • What never goes in a prompt. Secrets, customer data and anything under a data residency obligation. This matters most if your data must stay in Australia.
  • Least privilege for coding agents. Agents work in a sandbox or on a branch, with no production credentials, and only the tools and integrations you have approved.
  • A record. Pull requests note when AI did most of the work, so you can see where problems come from.
06 · Measurement

How we prove it’s working

We measure before we change anything, then again after, from your own repository and incident history:

MetricWhat it tells you
Change failure rateWhat share of releases need a fix or a rollback
Bugs found in productionHow many problems got past every check
Merges blocked by a checkHow many problems were caught before a reviewer saw them
Time from pull request to mergeWhether the checks are slowing the team down
07 · Delivery

How we deliver it

StageWhat happensYou get
AssessWe trace how recent production bugs got through, and review your pipeline, branch rules and AI tool useA short list of checks, in order of impact, with a price
BuildChecks, branch protection, code owners and the AI tool policy, added to your existing pipelineWorking guardrails on every pull request
TuneThresholds adjusted so the checks catch problems without blocking good workBefore and after measurements
Hand overRunbooks and a walkthrough for your developersEverything as code in your repositories
08 · FAQ

Common questions

Will this slow our developers down?

Not in practice. Checks run automatically in minutes, before anyone reviews the change. We tune thresholds so they block real problems, and we measure time to merge before and after to prove it.

Do you stop developers using AI?

No. We help you choose approved tools and set clear rules for them. The goal is to keep the speed and catch more of the mistakes.

Which AI tools does this work with?

All of them. The checks run on the code in the pull request, so it doesn’t matter whether it came from GitHub Copilot, Claude Code, Cursor or a person.

We use GitLab or Azure DevOps. Does that matter?

No. The same checks and review rules work in GitLab CI and Azure Pipelines, using their own branch protection and required-reviewer rules (code owners in GitLab, path-based reviewer policies in Azure Repos).

Next step

Show us your last incident

On a 30-minute call we’ll look at how the bug got through, and which one or two checks would have stopped it.