The short answer
Quick answer: Infrastructure as code (IaC) means defining your servers, networks, databases and other infrastructure in text files, and letting a tool create and update the real resources to match. Instead of clicking through a cloud console and hoping you remember every setting, you write down the desired result, keep it in version control, review changes like any other code, and apply them automatically. That makes infrastructure repeatable (the same every time), reviewable (changes are visible before they happen), recoverable (you can rebuild from the files) and documented (the files are the documentation).
What goes wrong with clicking
Building infrastructure by hand in a web console, sometimes called "ClickOps", is fine for learning and for one-off experiments. It breaks down quickly beyond that.
- Not repeatable. Creating a second identical environment means remembering dozens of settings.
- Drift. Staging and production gradually differ in small, undocumented ways, and bugs appear in one but not the other.
- No history. Who opened that firewall port, when, and why?
- No review. A mistake goes live the moment you click.
- Slow and error-prone. People make typos.
- Hard to recover. If an environment is lost, rebuilding it is archaeology.
- Snowflake servers. Each one has been hand-tuned until it is unique and nobody dares to replace it.
Martin Fowler's note on Infrastructure as Code describes the shift: treat infrastructure with the same practices as software.
Declarative: describe the result
Most IaC tools are declarative. You state what should exist, not the steps to create it.
resource "aws_s3_bucket" "assets" {
bucket = "example-app-assets"
}
resource "aws_instance" "web" {
count = 3
ami = "ami-0abcdef1234567890"
instance_type = "t3.small"
tags = {
Name = "web-${count.index}"
Environment = "production"
}
}
This says: one storage bucket and three servers of this type should exist. The tool works out how to get there.
| Imperative | Declarative | |
|---|---|---|
| You write | The steps: "create a server, then another" | The result: "there are three servers" |
| Running it twice | May create duplicates | Changes nothing the second time |
| Handles existing resources | You must check yourself | The tool compares and adjusts |
That second property is idempotence: applying the same definition repeatedly gives the same result. It is what makes automation safe.
The workflow
Using Terraform as the example; its introduction describes the same cycle of write, plan and apply.
- Write the definition files.
- Plan. The tool compares three things: what your files say, what it recorded last time, and what actually exists. It prints the exact changes it would make.
Plan: 1 to add, 1 to change, 0 to destroy.
+ aws_instance.web[2]
~ aws_instance.web[0]
instance_type: "t3.micro" -> "t3.small"
- Review. A teammate reads the plan in a pull request.
- Apply. The tool calls the cloud provider's APIs to make the changes, in the right order.
The plan is the most valuable part. You see the consequences before anything happens, including the alarming ones, such as a line saying a database will be destroyed and recreated.
State
To know what it manages, the tool keeps a state file: a record mapping each resource in your code to the real resource in the cloud.
State needs care:
- Store it remotely in shared, access-controlled storage, not on a laptop.
- Lock it during changes so two people cannot apply at once.
- Protect it. It often contains sensitive values.
- Do not edit it by hand.
Drift
Drift is when reality no longer matches the code, usually because someone changed something manually. The next plan will show the difference, and applying will revert the manual change. The cure is cultural as much as technical: make all changes through code.
The tools
| Tool | Purpose | Style |
|---|---|---|
| Terraform, OpenTofu | Provision cloud resources on any provider | Declarative, its own language (HCL) |
| AWS CloudFormation, Azure Bicep | Provision resources on one cloud | Declarative, YAML/JSON or a dedicated language |
| Pulumi, AWS CDK | Provision resources | Declarative results written in general-purpose languages |
| Ansible, Chef, Puppet | Configure software inside servers | Mostly declarative or procedural |
| Kubernetes manifests, Helm | Define workloads in a cluster | Declarative; see what Kubernetes does |
| Packer, Dockerfiles | Build machine and container images | Scripts |
Two categories are worth separating:
- Provisioning tools create infrastructure: networks, servers, databases.
- Configuration management tools set up what runs inside a server: packages, files, services.
Mutable vs immutable infrastructure
Mutable: servers are long-lived and updated in place. Over time each accumulates its own history of patches and fixes.
Immutable: servers are never changed after creation. To update, you build a new image, start new servers from it, and delete the old ones.
Immutable infrastructure removes drift between servers, makes rollback simple (start the old image again), and fits naturally with containers and with blue-green and canary deployments. The common phrase is to treat servers as cattle, not pets.
What you gain
- Version history. Every change is a commit with an author and a reason. See how Git works.
- Review before change.
- Reproducible environments. Create an identical staging environment, or a temporary one for each pull request, from the same code.
- Disaster recovery. Rebuild everything in a new region from the files.
- Automated checks. Policy tools can block insecure settings, such as publicly readable storage, before they are applied.
- Reuse. Package common patterns as modules: a standard network, a standard service.
- Speed. Creating an environment takes minutes.
- Living documentation. To know how production is set up, read the code.
GitOps
GitOps takes this one step further: the Git repository is the single source of truth, and an automated agent continuously makes the running system match it. Nobody runs commands against production directly. A change is a merged pull request, and a rollback is a revert. Tools such as Argo CD and Flux do this for Kubernetes.
Pitfalls
- Secrets in code or state. Use a secrets manager and reference values from it.
- Huge monolithic configurations. One mistake can affect everything, and plans become slow. Split by environment and by component to limit the blast radius.
- Manual changes during an incident that are never brought back into the code.
- Not reading the plan. "Replace" on a database means destroy and recreate.
- Unpinned versions of providers and modules, so behaviour changes unexpectedly.
- Copy and paste between environments in place of modules and variables.
- Importing existing infrastructure is tedious, though tools exist to help.
Good practice
- Keep infrastructure code in version control, with review required.
- Run plan and apply from a CI/CD pipeline, not from laptops.
- Use remote state with locking.
- Separate environments, and use the same modules for each.
- Pin versions.
- Tag resources so their purpose and owner are clear.
- Add automated policy and security checks.
- Protect critical resources from accidental deletion.
Frequently asked questions
What is infrastructure as code in simple terms?
Describing your servers, networks and other infrastructure in text files, and using a tool to create and update the real thing to match those files.
What is the difference between Terraform and Ansible?
Terraform is mainly for provisioning infrastructure resources. Ansible is mainly for configuring the software on servers. They are often used together.
What is configuration drift?
The gap that opens up when the real infrastructure no longer matches its definition, usually because of manual changes.
What is Terraform state?
A record the tool keeps that maps the resources in your code to the actual resources it created, so it knows what to change next time.
Conclusion
Clicking around a console creates infrastructure that exists only in the cloud and in someone's memory. Infrastructure as code puts it in files that can be versioned, reviewed, tested and reapplied. The habit to build is simple and strict: if it is not in code, it does not get changed.
Related articles
- How CI/CD Pipelines Ship Code Automatically
- What Kubernetes Does and Why Companies Use It
- How Git Stores Your Code History
- How Auto-Scaling Handles Sudden Traffic Spikes
