Where the time actually goes
Slow releases are rarely slow builds. When we trace a change from commit to production, most of the time is spent waiting:
- Waiting for a shared test environment that someone else is using, or that drifted from production.
- Waiting for a release window, because deploying is manual and risky, so it is batched.
- Waiting on flaky tests that get re-run until they pass.
- Waiting for the one person who knows how to deploy.
Batching makes it worse. The less often you release, the more each release contains, the riskier it is, and the more carefully it has to be scheduled. The fix is to make each release small, automatic and reversible.
The reference pipeline
Every change to the main branch runs the same path. No stage needs a person unless your policy requires an approval.
| Stage | What it does | Fails the release when |
|---|---|---|
| Build | Builds one immutable container image, tagged with the commit SHA | The build breaks |
| Test | Unit and integration tests in parallel | Any test fails; flaky tests are quarantined, not retried |
| Scan | Dependency, secret and container image scanning | A critical vulnerability or committed secret is found |
| Publish | Pushes the image to ECR or ACR | n/a |
| Staging | Deploys the same image to staging and runs smoke tests | Smoke tests fail |
| Production | Blue/green or canary release, watched by alarms | Error rate or latency alarms fire; traffic shifts back automatically |
The same image moves through every environment. Nothing is rebuilt for production, so what you tested is what you ship.
What it looks like in code
A trimmed GitHub Actions workflow for an ECS service. It signs in to AWS with OIDC, so there are no long-lived access keys stored in GitHub.
name: deploy on: push: branches: [main] permissions: id-token: write # lets the job request an OIDC token from AWS contents: read jobs: release: runs-on: ubuntu-latest environment: production steps: - uses: actions/checkout@v4 - name: Test run: make test - uses: aws-actions/configure-aws-credentials@v4 with: role-to-assume: arn:aws:iam::123456789012:role/github-deploy aws-region: ap-southeast-2 - id: ecr uses: aws-actions/amazon-ecr-login@v2 - id: image name: Build and push run: | IMAGE=${{ steps.ecr.outputs.registry }}/app:${{ github.sha }} docker build -t "$IMAGE" . docker push "$IMAGE" echo "uri=$IMAGE" >> "$GITHUB_OUTPUT" - id: taskdef uses: aws-actions/amazon-ecs-render-task-definition@v1 with: task-definition: ecs/task-definition.json container-name: app image: ${{ steps.image.outputs.uri }} - name: Blue/green deploy uses: aws-actions/amazon-ecs-deploy-task-definition@v2 with: task-definition: ${{ steps.taskdef.outputs.task-definition }} cluster: prod service: app codedeploy-appspec: ecs/appspec.yaml codedeploy-application: app codedeploy-deployment-group: app-prod wait-for-service-stability: true
The deploy role can do one thing.
The IAM role GitHub assumes is scoped to this repository and branch, and can only update this service. A leaked workflow can’t touch anything else in the account.
Infrastructure as code
Every environment is built from the same Terraform modules with different inputs, so staging is a smaller copy of production rather than a close guess. Changes go through pull requests like application code.
- Plan on pull request. The planned change is posted on the PR for review.
- Apply on merge. Only the pipeline can apply; people don’t have write access to production by default.
- Drift detection. A scheduled plan flags any change made by hand in the console.
terraform { backend "s3" { bucket = "acme-terraform-state" key = "prod/app.tfstate" region = "ap-southeast-2" encrypt = true use_lockfile = true # S3-native state locking, Terraform 1.10+ } }
One path for app and infrastructure changes
Slow releases are often a handoff problem, not a tooling one. A developer needs a queue or a database, raises a ticket, and waits for an infrastructure engineer to get to it. We close that gap in three ways:
- Same repository, same review. Infrastructure changes ride in the same pull request as the code that needs them, and go through the same pipeline.
- Modules developers can use. Common building blocks (a queue, a database, a bucket) are reusable modules with sensible defaults, so asking for infrastructure is a few lines of code.
- Platform engineers review, not gatekeep. They own the modules and the guardrails, and approve changes on the pull request instead of doing them by hand.
We wrote about the people side of this in Closing the gap between Dev and DevOps.
Releases you can take back
A release is safe when rolling it back is faster than debugging it. We pick the strategy per service.
| Strategy | How it works | Rollback | Tooling |
|---|---|---|---|
| Blue/green | New version runs beside the old one; traffic switches when it is healthy | Switch traffic back | ECS + CodeDeploy, App Service deployment slots |
| Canary | A small share of traffic goes to the new version first, then more as alarms stay quiet | Automatic on alarm | Argo Rollouts on EKS, Lambda aliases with CodeDeploy |
| Rolling | Instances are replaced a few at a time | Roll forward or redeploy previous image | ECS and Kubernetes defaults |
| Feature flags | Code ships switched off and is turned on separately | Turn the flag off | AppConfig, LaunchDarkly or your own |
Database changes follow expand and contract: add the new column or table first, ship the code that uses it, and remove the old one in a later release. Every release stays compatible with the one before it.
Security built into the path
- No long-lived cloud keys in CI. OIDC federation to AWS and Azure.
- Branch protection with required reviews and passing checks before merge.
- Secrets in AWS Secrets Manager or Azure Key Vault, never in the repository.
- Images scanned on push, with critical findings failing the build.
- Every production change is traceable to a commit, a pull request and a reviewer.
How we prove it got faster
We baseline the four DORA metrics in the first week and report them again at handover, from your own pipeline data.
| Metric | Question it answers |
|---|---|
| Deployment frequency | How often do you ship to production? |
| Change lead time | How long from commit to running in production? |
| Change fail rate | What share of releases need a fix or rollback? |
| Failed deployment recovery time | How long to recover when a release goes wrong? |
Speed and stability move together. Smaller, automated releases fail less often and recover faster.
Common questions
We use GitLab, Bitbucket or Azure DevOps. Does that matter?
No. The stages and controls are the same. We build on the tool you already have rather than migrating you.
Do we have to move to Kubernetes?
No. ECS Fargate or Azure Container Apps suit most teams and are far less to operate. We recommend EKS or AKS only when you need what they offer.
Can we keep a manual approval before production?
Yes. An approval gate is one setting on the production environment. It stays a click, not a runbook.
What about our existing environments built by hand?
We import them into Terraform where it is sensible and rebuild where it isn’t, one environment at a time, starting with the least risky.
Next step
Show us your last release
Walk us through how your last change reached production. On a 30-minute call we’ll point to where the time went.