Terraform Plan Shows Changes Every Time: Fixing the Perpetual Diff
Terraform · Intermediate · 6 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Your terraform plan reports the same 3 changes on every run, and applying them changes nothing. Here's how to find the normalization mismatch causing it.
The Problem
Terraform plan shows changes every time you run it. Same three resources, same attributes, on a branch where you edited nothing but a comment. You apply — it succeeds in 12 seconds — and the very next terraform plan reports the identical three changes again.
The blast radius is bigger than annoyance. Your "plan must be empty" CI gate is now permanently red, so someone disables it. Reviewers stop reading plan output because 90% of it is noise. Three weeks later a real deletion of a production RDS instance sails through review inside that noise.
Typical offenders:
~ resource "aws_iam_role_policy" "app" {
~ policy = jsonencode(
~ {
~ Statement = [ ... ]
}
)
}
~ resource "aws_ecs_task_definition" "api" {
~ container_definitions = (sensitive value)
}
Why the Obvious Fix Falls Short
The reflex is lifecycle { ignore_changes = [...] }. It makes the plan clean in one commit and it is almost always the wrong move.
ignore_changes operates at attribute granularity. Ignoring container_definitions on a task definition means the image tag lives in that attribute too — you have just told Terraform to stop deploying your application. Ignoring policy on an IAM policy means a teammate can widen it to Action: "*" in the console and Terraform will cheerfully report no changes forever.
The second reflex — "it must be console drift, let's lock down IAM" — also misses. Run terraform plan -refresh=false: no API call is made at all, and the diff is still there. That single command disproves the drift theory in five seconds. A diff that survives with refresh disabled is a comparison between your rendered config and your state file. Nothing external is involved.
The third reflex, bumping the provider version, occasionally helps but usually just changes which attribute perpetually diffs.
How It Actually Works
Terraform's plan is a three-way comparison: prior state, refreshed remote value, and your rendered config. A perpetual diff means one specific edge: the value you write is not byte-identical to the value the provider persists after a round trip through the API.
Cloud APIs normalize. AWS canonicalizes IAM policy JSON — reorders keys, collapses whitespace, may rewrite a single-element Action string into an array. ECS injects defaults you never specified ("cpu": 0, "essential": true, "mountPoints": []). ARNs come back with different casing. A heredoc block carries a trailing newline that the API strips.
Terraform compares strings. {"Effect":"Allow","Action":"s3:GetObject"} and {"Action":["s3:GetObject"],"Effect":"Allow"} are semantically identical and textually different, so you get a diff — forever.
The other family of causes is config that isn't deterministic: timestamp(), uuid(), a random_* resource without keepers, or a data source whose result changes on every read.
flowchart TD
A["HCL config renders\npolicy = heredoc JSON"] --> B["terraform apply\nsends string to API"]
B --> C["AWS canonicalizes:\nreorders keys, adds defaults"]
C --> D["Provider reads value back\nwrites canonical form to state"]
D --> E{"Next plan:\nrendered config ==\nstate string?"}
E -->|"byte-identical"| F["No changes"]
E -->|"differs in form only"| G["Perpetual diff\napply is a no-op"]
G --> B
Note the loop: apply feeds the un-normalized string back, AWS re-normalizes, state gets the canonical form again. Nothing ever converges.
Before and After
# BEFORE — heredoc JSON. Whitespace, key order and Action-as-string
# all get rewritten by AWS, so state never matches config.
resource "aws_iam_role_policy" "app" {
role = aws_iam_role.app.id
policy = <<-EOT
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": "s3:GetObject",
"Resource": "${aws_s3_bucket.data.arn}/*"
}]
}
EOT
}
resource "aws_instance" "worker" {
ami = data.aws_ami.latest.id
instance_type = "t3.medium"
# timestamp() re-evaluates on EVERY plan -> forced replacement, every run
tags = { LastDeployed = timestamp() }
}
# AFTER — let the provider produce the canonical string, and remove
# non-deterministic values from config entirely.
data "aws_iam_policy_document" "app" {
statement {
effect = "Allow"
actions = ["s3:GetObject"] # already a list, like AWS returns
resources = ["${aws_s3_bucket.data.arn}/*"]
}
}
resource "aws_iam_role_policy" "app" {
role = aws_iam_role.app.id
# .json is emitted in the same canonical shape AWS stores
policy = data.aws_iam_policy_document.app.json
}
resource "aws_instance" "worker" {
ami = data.aws_ami.latest.id
instance_type = "t3.medium"
# Deterministic: only changes when the thing being deployed changes
tags = { DeployedRevision = var.git_sha }
}
For container_definitions, the equivalent fix is jsonencode([...]) with the AWS-injected defaults written out explicitly (essential, cpu, mountPoints, volumesFrom, portMappings[].protocol). Copy them straight out of terraform state show.
When NOT to Use This
- Genuinely externally-managed attributes. If an autoscaler owns
desired_countor an external secret rotator owns a password,ignore_changeson that one narrow attribute is correct — that's what it's for. The rule is: ignore things another system legitimately owns, never things you own but encoded badly. - Confirmed provider bugs. Some attributes are unrepresentable (write-only fields returned as
null, or""vs unset ambiguity). Useignore_changeswith a comment linking the GitHub issue and a date to revisit. data.aws_iam_policy_documentisn't universal. For non-AWS providers or policies withNotAction/condition edge cases, plainjsonencode(...)with a hand-matched key set is often the pragmatic route.- Deleting the attribute isn't the answer either. Removing
tagsbecause it diffs just means the tag silently disappears on next apply.
Gotchas
tagsvstags_all. With provider-leveldefault_tags, setting the same key in a resource'stagsproduces a permanent diff in some provider versions. Pick one place to define each key.- Sets vs lists. Security group rules and similar attributes are sets; if the provider models them as lists, reordering in your config or in the API response creates a phantom diff. Sorting your config to match the API response is a band-aid — check whether a nested block form exists.
ignore_changesdoesn't apply at create time. The initial apply still uses your config value, so a wrong value gets baked in and then permanently hidden.- Empty string vs null.
description = ""and omittingdescriptionare different in Terraform but often identical to the API. This is the single most common one-attribute perpetual diff. -refresh=falsein CI hides real drift. It's a great diagnostic, a bad default — you'll stop noticing manual console changes.terraform planon a module you don't own may diff because the module hardcodes a value your provider defaults differently. Checkterraform state show <addr>before blaming yourself.
Key takeaway: If `apply` then immediately `plan` still shows the diff, it isn't drift — your config text and the value the API returns aren't byte-identical, so fix the encoding instead of reaching for `ignore_changes`.
Real-world challenge
Your CI pipeline gates merges on `terraform plan` producing no changes. For the past week every plan on `main` shows the same two changes on an `aws_ecs_task_definition` and an `aws_iam_role_policy`, even on commits that touched only README files. Applying succeeds, then the next plan shows the same two changes again. No one has console access. How do you diagnose and fix it?
1. Prove it's not real drift. Apply, then immediately:
terraform plan -refresh=false -out=tf.plan
terraform show -json tf.plan | jq '.resource_changes[] | select(.change.actions != ["no-op"]) | {addr:.address, before:.change.before, after:.change.after}'
If the diff survives with refresh disabled, nothing drifted remotely — config text ≠ state text.
2. Diff the two values character by character. Dump both sides to files and diff them. For the IAM policy you'll typically see the same statements with reordered keys and collapsed whitespace: AWS canonicalizes policy JSON. For the task definition you'll see fields you never set (cpu: 0, essential: true, empty mountPoints: []) that AWS injects.
3. Fix the encoding, not the symptom.
# IAM: let the provider canonicalize, not your heredoc
data "aws_iam_policy_document" "app" { statement { ... } }
resource "aws_iam_role_policy" "app" { policy = data.aws_iam_policy_document.app.json }
# ECS: jsonencode + spell out the defaults AWS adds back
container_definitions = jsonencode([{ name = "app", essential = true, cpu = 0, mountPoints = [], volumesFrom = [], ... }])
4. Re-run. Plan should be clean. Only if a provider bug makes the value genuinely unrepresentable do you fall back to a narrowly scoped lifecycle { ignore_changes = [container_definitions] } — and add a comment linking the provider issue so it gets removed.