Terraform Plan Shows Changes Every Time: Fixing the Perpetual Diff

Terraform · Intermediate · 6 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Your terraform plan reports the same 3 changes on every run, and applying them changes nothing. Here's how to find the normalization mismatch causing it.

The Problem

Terraform plan shows changes every time you run it. Same three resources, same attributes, on a branch where you edited nothing but a comment. You apply — it succeeds in 12 seconds — and the very next terraform plan reports the identical three changes again.

The blast radius is bigger than annoyance. Your "plan must be empty" CI gate is now permanently red, so someone disables it. Reviewers stop reading plan output because 90% of it is noise. Three weeks later a real deletion of a production RDS instance sails through review inside that noise.

Typical offenders:

~ resource "aws_iam_role_policy" "app" {
    ~ policy = jsonencode(
        ~ {
            ~ Statement = [ ... ]
          }
      )
  }
~ resource "aws_ecs_task_definition" "api" {
    ~ container_definitions = (sensitive value)
  }

Why the Obvious Fix Falls Short

The reflex is lifecycle { ignore_changes = [...] }. It makes the plan clean in one commit and it is almost always the wrong move.

ignore_changes operates at attribute granularity. Ignoring container_definitions on a task definition means the image tag lives in that attribute too — you have just told Terraform to stop deploying your application. Ignoring policy on an IAM policy means a teammate can widen it to Action: "*" in the console and Terraform will cheerfully report no changes forever.

The second reflex — "it must be console drift, let's lock down IAM" — also misses. Run terraform plan -refresh=false: no API call is made at all, and the diff is still there. That single command disproves the drift theory in five seconds. A diff that survives with refresh disabled is a comparison between your rendered config and your state file. Nothing external is involved.

The third reflex, bumping the provider version, occasionally helps but usually just changes which attribute perpetually diffs.

How It Actually Works

Terraform's plan is a three-way comparison: prior state, refreshed remote value, and your rendered config. A perpetual diff means one specific edge: the value you write is not byte-identical to the value the provider persists after a round trip through the API.

Cloud APIs normalize. AWS canonicalizes IAM policy JSON — reorders keys, collapses whitespace, may rewrite a single-element Action string into an array. ECS injects defaults you never specified ("cpu": 0, "essential": true, "mountPoints": []). ARNs come back with different casing. A heredoc block carries a trailing newline that the API strips.

Terraform compares strings. {"Effect":"Allow","Action":"s3:GetObject"} and {"Action":["s3:GetObject"],"Effect":"Allow"} are semantically identical and textually different, so you get a diff — forever.

The other family of causes is config that isn't deterministic: timestamp(), uuid(), a random_* resource without keepers, or a data source whose result changes on every read.

flowchart TD
    A["HCL config renders\npolicy = heredoc JSON"] --> B["terraform apply\nsends string to API"]
    B --> C["AWS canonicalizes:\nreorders keys, adds defaults"]
    C --> D["Provider reads value back\nwrites canonical form to state"]
    D --> E{"Next plan:\nrendered config ==\nstate string?"}
    E -->|"byte-identical"| F["No changes"]
    E -->|"differs in form only"| G["Perpetual diff\napply is a no-op"]
    G --> B

Note the loop: apply feeds the un-normalized string back, AWS re-normalizes, state gets the canonical form again. Nothing ever converges.

Before and After

# BEFORE — heredoc JSON. Whitespace, key order and Action-as-string
# all get rewritten by AWS, so state never matches config.
resource "aws_iam_role_policy" "app" {
  role   = aws_iam_role.app.id
  policy = <<-EOT
    {
      "Version": "2012-10-17",
      "Statement": [{
        "Effect": "Allow",
        "Action": "s3:GetObject",
        "Resource": "${aws_s3_bucket.data.arn}/*"
      }]
    }
  EOT
}

resource "aws_instance" "worker" {
  ami           = data.aws_ami.latest.id
  instance_type = "t3.medium"
  # timestamp() re-evaluates on EVERY plan -> forced replacement, every run
  tags = { LastDeployed = timestamp() }
}
# AFTER — let the provider produce the canonical string, and remove
# non-deterministic values from config entirely.
data "aws_iam_policy_document" "app" {
  statement {
    effect    = "Allow"
    actions   = ["s3:GetObject"]          # already a list, like AWS returns
    resources = ["${aws_s3_bucket.data.arn}/*"]
  }
}

resource "aws_iam_role_policy" "app" {
  role = aws_iam_role.app.id
  # .json is emitted in the same canonical shape AWS stores
  policy = data.aws_iam_policy_document.app.json
}

resource "aws_instance" "worker" {
  ami           = data.aws_ami.latest.id
  instance_type = "t3.medium"
  # Deterministic: only changes when the thing being deployed changes
  tags = { DeployedRevision = var.git_sha }
}

For container_definitions, the equivalent fix is jsonencode([...]) with the AWS-injected defaults written out explicitly (essential, cpu, mountPoints, volumesFrom, portMappings[].protocol). Copy them straight out of terraform state show.

When NOT to Use This

Gotchas

Key takeaway: If `apply` then immediately `plan` still shows the diff, it isn't drift — your config text and the value the API returns aren't byte-identical, so fix the encoding instead of reaching for `ignore_changes`.

Real-world challenge

Your CI pipeline gates merges on `terraform plan` producing no changes. For the past week every plan on `main` shows the same two changes on an `aws_ecs_task_definition` and an `aws_iam_role_policy`, even on commits that touched only README files. Applying succeeds, then the next plan shows the same two changes again. No one has console access. How do you diagnose and fix it?

1. Prove it's not real drift. Apply, then immediately:

terraform plan -refresh=false -out=tf.plan
terraform show -json tf.plan | jq '.resource_changes[] | select(.change.actions != ["no-op"]) | {addr:.address, before:.change.before, after:.change.after}'

If the diff survives with refresh disabled, nothing drifted remotely — config text ≠ state text.

2. Diff the two values character by character. Dump both sides to files and diff them. For the IAM policy you'll typically see the same statements with reordered keys and collapsed whitespace: AWS canonicalizes policy JSON. For the task definition you'll see fields you never set (cpu: 0, essential: true, empty mountPoints: []) that AWS injects.

3. Fix the encoding, not the symptom.

# IAM: let the provider canonicalize, not your heredoc
data "aws_iam_policy_document" "app" { statement { ... } }
resource "aws_iam_role_policy" "app" { policy = data.aws_iam_policy_document.app.json }

# ECS: jsonencode + spell out the defaults AWS adds back
container_definitions = jsonencode([{ name = "app", essential = true, cpu = 0, mountPoints = [], volumesFrom = [], ... }])

4. Re-run. Plan should be clean. Only if a provider bug makes the value genuinely unrepresentable do you fall back to a narrowly scoped lifecycle { ignore_changes = [container_definitions] } — and add a comment linking the provider issue so it gets removed.