How Terraform goes wrong as a platform grows
Most Terraform starts as one directory that builds everything. That works for the first few services. Then every new service is copied from the last one with small changes, the plan for a one-line change takes minutes and touches hundreds of resources, and nobody wants to run apply.
Anwar has designed a reusable Terraform module structure used to rebuild a platform on one consistent pattern. This is the approach in outline.
Two layers: modules and stacks
- Modules describe one building block: a network, a service, a database, a queue. They take inputs, create resources and return outputs. They don’t know which environment they’re in.
- Stacks (root modules) combine modules for one part of one environment, such as “production networking” or “staging, orders service”. Each stack has its own state.
Every environment uses the same modules with different inputs. Staging becomes a smaller copy of production rather than a close guess.
What makes a good module
- One job. A “service” module that creates the container service, its load balancer rule, log group, alarms and IAM role is easier to reuse than one that also creates the network.
- Few, meaningful inputs with safe defaults. Encryption on, logs retained, deletion protection on in production. A new service should need a handful of values, not fifty.
- The standard layout.
main.tf,variables.tf,outputs.tfand a README, so anyone can read any module. - No provider configuration inside. Child modules inherit providers from the stack that calls them.
for_eachovercountfor collections, so adding or removing one item doesn’t renumber and recreate the others.
Version modules, and upgrade one stack at a time
Publish modules with version tags, from a registry or a Git tag, and pin each stack to a version. A change to a module then reaches staging first, and production only when someone moves its pin. Without versions, changing a module changes every environment at the next apply.
Refactoring without recreating resources
Moving existing resources into modules changes their addresses, and Terraform would otherwise destroy and recreate them. A moved block tells Terraform it’s the same resource:
moved {
from = aws_s3_bucket.logs
to = module.logging.aws_s3_bucket.this
}For infrastructure that was built in the console, an import block brings it under Terraform in a normal plan and apply, so the change can be reviewed like any other.
State: small, separate and locked
Keep state remote, in S3 with locking, with one state file per stack. Smaller states mean faster plans, and a mistake in one stack can’t reach another. Read another stack’s outputs through data sources or remote state, rather than sharing one big state.
Run it in the pipeline
Plan on every pull request and post the plan for review. Apply only from the pipeline after merge. Schedule a plan to flag anything changed by hand. See Terraform and infrastructure as code and Faster deployments.
Sources
Next step
Want a review of your Terraform?
On a 30-minute call we’ll look at how your code is structured and where to start improving it.