Docker build cache not working in GitHub Actions: exporting layers that actually get reused

CI/CD · Intermediate · 6 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Your Docker build takes 40 seconds locally and 6 minutes on every CI run because the runner starts with an empty layer store. Here's how to export and restore layers correctly.

The Problem

Docker build cache not working in GitHub Actions is one of those failures with no error message at all — the build just succeeds, slowly, forever. Locally docker build finishes in 12 seconds because nothing changed; the same Dockerfile on a runner spends 5m40s re-running apt-get install, npm ci, and a Rust compile on every single push, including pushes that only touched a README.

The build log gives it away: not one CACHED line. Every RUN executes from scratch.

#8 [deps 2/2] RUN npm ci
#8 DONE 214.3s

On a busy repo that's 40+ wasted minutes a day and a feedback loop long enough that people stop waiting for CI.

Why the Obvious Fix Falls Short

The instinct is to pull the previous image and use it as a cache source:

- run: docker pull ghcr.io/org/app:latest || true
- run: docker build --cache-from ghcr.io/org/app:latest -t app .

This mostly doesn't work, for a reason that isn't obvious: a normal pushed image contains only the final filesystem layers plus a config. It carries no record of which build step produced which layer. Without that metadata BuildKit can't match your RUN npm ci against anything in the pulled image, so it rebuilds.

Even when you add BUILDKIT_INLINE_CACHE=1 to embed that metadata, you've only fixed the last stage. In a multi-stage Dockerfile — the standard layout, where a fat builder stage compiles and a slim runtime stage copies the artifact — the expensive stages aren't in the exported image at all. Inline cache gives you hits on the three cheap COPY lines and misses on the 3-minute compile.

The second instinct is actions/cache on /var/lib/docker. That directory isn't reliably readable/writable around a running daemon, restores take minutes because it's enormous, and with the containerd snapshotter the layout changed anyway. People who try this usually end up with a 4 GB tarball that takes longer to restore than the build took.

How It Actually Works

The core fact: a GitHub-hosted runner is a fresh VM. There is no persistent layer store. Caching isn't "turning on" a cache that already exists — it's exporting your layer graph to external storage at the end of the build and importing it at the start of the next one.

BuildKit does this through cache exporters. docker/setup-buildx-action gives you a BuildKit builder that supports:

Two knobs decide whether it actually helps:

flowchart TD
    A[push to branch] --> B[setup-buildx: fresh BuildKit, empty store]
    B --> C{cache-from: import<br/>manifest + layer blobs}
    C -->|hit| D[match step digests<br/>against imported graph]
    C -->|miss| E[execute every RUN]
    D -->|digest matches| F[CACHED - pull layer blob]
    D -->|digest differs| G[execute step +<br/>all steps after it]
    F --> H[final image]
    G --> H
    E --> H
    H --> I[cache-to mode=max:<br/>export ALL stage layers]
    I --> J[(GHA cache / registry<br/>scoped by branch)]
    J -.next run.-> C

Note the digest differs branch: cache matching is a prefix match over the step chain. Once one step's inputs change, everything after it rebuilds regardless of what's in the cache. That's why COPY . . before npm ci destroys cache effectiveness even when caching is configured perfectly.

Before and After

# BEFORE - plain docker build on a fresh VM: zero reuse, always cold
jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: docker pull ghcr.io/org/app:latest || true
      # pulled image has no per-step cache metadata, and the builder
      # stage isn't in it at all -> every RUN re-executes
      - run: |
          docker build --cache-from ghcr.io/org/app:latest \
            -t ghcr.io/org/app:${{ github.sha }} .
# BEFORE - copies everything before installing, so any source edit
# invalidates the dependency layer
FROM node:22-slim AS deps
WORKDIR /app
COPY . .
RUN npm ci
# AFTER - BuildKit builder + explicit export/import, mode=max, stable scope
jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: docker/setup-buildx-action@v3   # required: gives a BuildKit builder
      - uses: docker/build-push-action@v6
        with:
          push: true
          tags: ghcr.io/org/app:${{ github.sha }}
          # mode=max exports intermediate stages too - without it the
          # expensive builder stage is never cached
          cache-from: type=gha,scope=${{ github.workflow }}
          cache-to: type=gha,mode=max,scope=${{ github.workflow }}
# AFTER - lockfile-only copy so the install layer survives source edits
FROM node:22-slim AS deps
WORKDIR /app
COPY package.json package-lock.json ./
RUN --mount=type=cache,target=/root/.npm npm ci
COPY . .

Important pairing: RUN --mount=type=cache gives you a package-manager cache within a build, but that mount is not exported by type=gha. It only helps when the layer above it is invalidated but BuildKit still has the mount from the same builder. On ephemeral runners, the layer cache is what saves you; treat the cache mount as a bonus for self-hosted runners.

When NOT to Use This

Gotchas

Key takeaway: A GitHub runner has no layer store to reuse, so caching only works if you explicitly export layers with cache-to mode=max and scope them so branch builds can actually read them.

Real-world challenge

A team added `cache-from`/`cache-to: type=gha,mode=max` and saw builds drop from 7 minutes to 50 seconds — for about two weeks. Now cache hits are erratic: main branch builds hit, but PR builds almost always miss, and even main misses after the nightly multi-arch build job runs. Diagnose it.

Two separate causes, both about cache scope and size.

  1. PR misses: GitHub Actions cache isolates entries by branch. A PR branch can read caches written by its base branch, but two PRs can never see each other's. If the workflow only writes cache on PR events, each PR writes into its own isolated scope and nobody reuses it. Fix: run the build on pushes to main too, so a warm base-branch cache exists for PRs to read from.

  2. Main misses after the nightly job: the repo-wide Actions cache is capped at 10 GB with LRU eviction. A mode=max multi-arch build exports two full layer sets and can easily be several GB, evicting everything else. Give it a separate, bounded scope — or move the big one to a registry cache which isn't subject to the 10 GB budget.

- uses: docker/build-push-action@v6
  with:
    cache-from: type=gha,scope=${{ github.workflow }}-amd64
    cache-to: type=gha,mode=max,scope=${{ github.workflow }}-amd64

For the nightly multi-arch job, switch to type=registry,ref=ghcr.io/org/app:buildcache,mode=max so it stops competing for the 10 GB budget.

Verify by grepping the build log for CACHED step counts and by checking Cache size in the Actions caches UI after each run.