Code is written faster than it is checked
AI coding assistants and agents let one developer produce far more code in a day. Review time, test suites and release processes didn’t speed up to match. Pull requests get bigger, reviewers skim, and the code looks right, which is exactly what makes its mistakes easy to miss.
AI also brings new kinds of mistakes:
- Invented dependencies. A suggested package or API that doesn’t exist, or a look-alike name that an attacker can register.
- Tests that prove nothing. Tests written to pass against the code, rather than to check what it should do.
- Confident edge-case bugs. The happy path works, and the error handling, time zones and permissions don’t.
- Data in prompts. Secrets, customer records or internal code pasted into tools that were never approved.
We build with AI every day at DevOpsPlant, and we govern it. The answer isn’t to slow developers down or ban the tools. It is to make every change prove itself automatically, and keep humans on the decisions that matter.
The checks every pull request must pass
These run on every pull request, however the code was written. A developer gets the result in minutes, before a reviewer spends any time on it.
| Check | What it catches | Blocks the merge when |
|---|---|---|
| Tests on changed lines | New code with no tests, or tests that don’t exercise it | Coverage of the changed lines is below the agreed minimum |
| Security scan | Injection, unsafe deserialisation and other common flaws | A new finding at high severity or above |
| Dependency check | Packages that don’t exist, look-alike names, known vulnerabilities, unapproved licences | A new package isn’t on the approved list |
| Secret scanning | Keys, tokens and passwords in the diff | Any secret is found |
| Change size | Pull requests too large to review properly | The diff passes an agreed size, unless a lead approves the exception |
Coverage of the changed lines matters more than overall coverage. A codebase can sit at 70% while the new, AI-written code has no tests at all.
# Fail the pull request if the lines it changes are under 80% covered - name: Coverage of changed lines run: | pip install diff-cover diff-cover coverage.xml --compare-branch=origin/main --fail-under=80
Humans still approve
AI review tools are useful as a first pass, flagging obvious problems before a person looks. They don’t approve changes.
- Every merge needs a human approval. Branch protection enforces it, including for changes an agent opens.
- Sensitive code has named owners. Payments, authentication, infrastructure and the pipeline itself need a reviewer from the owning team.
- Small pull requests by default. A reviewer can properly check a few hundred lines, not a few thousand.
# Changes here always need a reviewer from the owning team, however the code was written
/src/payments/ @acme/payments-leads
/src/auth/ @acme/security
/infra/ @acme/platform
/.github/ @acme/platformAssume some bugs will still get through
No set of checks catches everything, so releases are built to limit the damage:
- Feature flags. New features ship switched off, then turn on for a small group first.
- Canary releases with automatic rollback. A new version takes a small share of traffic, and rolls back by itself if error rates rise.
- Fast rollback. Going back is one step, practised, and faster than debugging in production.
The release side is covered in detail in From commit to production in minutes.
Rules for the AI tools themselves
Governance is a short, written policy your developers can follow, enforced where it can be:
- Approved tools. A list of the assistants and agents the team may use, on business accounts with training on your code turned off.
- What never goes in a prompt. Secrets, customer data and anything under a data residency obligation. This matters most if your data must stay in Australia.
- Least privilege for coding agents. Agents work in a sandbox or on a branch, with no production credentials, and only the tools and integrations you have approved.
- A record. Pull requests note when AI did most of the work, so you can see where problems come from.
How we prove it’s working
We measure before we change anything, then again after, from your own repository and incident history:
| Metric | What it tells you |
|---|---|
| Change failure rate | What share of releases need a fix or a rollback |
| Bugs found in production | How many problems got past every check |
| Merges blocked by a check | How many problems were caught before a reviewer saw them |
| Time from pull request to merge | Whether the checks are slowing the team down |
How we deliver it
| Stage | What happens | You get |
|---|---|---|
| Assess | We trace how recent production bugs got through, and review your pipeline, branch rules and AI tool use | A short list of checks, in order of impact, with a price |
| Build | Checks, branch protection, code owners and the AI tool policy, added to your existing pipeline | Working guardrails on every pull request |
| Tune | Thresholds adjusted so the checks catch problems without blocking good work | Before and after measurements |
| Hand over | Runbooks and a walkthrough for your developers | Everything as code in your repositories |
Common questions
Will this slow our developers down?
Not in practice. Checks run automatically in minutes, before anyone reviews the change. We tune thresholds so they block real problems, and we measure time to merge before and after to prove it.
Do you stop developers using AI?
No. We help you choose approved tools and set clear rules for them. The goal is to keep the speed and catch more of the mistakes.
Which AI tools does this work with?
All of them. The checks run on the code in the pull request, so it doesn’t matter whether it came from GitHub Copilot, Claude Code, Cursor or a person.
We use GitLab or Azure DevOps. Does that matter?
No. The same checks and review rules work in GitLab CI and Azure Pipelines, using their own branch protection and required-reviewer rules (code owners in GitLab, path-based reviewer policies in Azure Repos).
Next step
Show us your last incident
On a 30-minute call we’ll look at how the bug got through, and which one or two checks would have stopped it.