Kubernetes Job Stuck Because the Sidecar Keeps Running: Native Sidecars Fix It

Kubernetes · Intermediate · 6 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Your Job's main container exits but the Pod stays Running forever because a proxy or log-shipper sidecar never stops. Native sidecar containers end that.

What Changed

If your Kubernetes Job is stuck because a sidecar keeps running, that's no longer something you have to hack around. Kubernetes now supports native sidecar containers: an entry in initContainers with restartPolicy: Always. Such a container starts before the regular containers, stays alive alongside them, is restarted independently if it crashes, and — the part that matters here — is terminated automatically by the kubelet once all regular containers exit, so the Pod can reach Succeeded and the Job can complete.

Availability: alpha in 1.28, beta and enabled by default in 1.29, GA in 1.33. If your cluster is 1.29 or newer you can use it today without a feature gate.

The Old Way vs The New Way

Before — a sentinel file so the sidecar knows when to give up:

apiVersion: batch/v1
kind: Job
metadata: { name: nightly-export }
spec:
  template:
    spec:
      restartPolicy: Never
      volumes:
        - name: signal
          emptyDir: {}
      containers:
        - name: export
          image: myapp/export:2.1
          volumeMounts: [{ name: signal, mountPath: /signal }]
          command: ["/bin/sh", "-c"]
          args:
            - |
              /app/export --db=127.0.0.1:5432
              code=$?
              touch /signal/done      # easy to forget on error paths
              exit $code
        - name: sql-proxy
          image: gcr.io/cloud-sql-connectors/cloud-sql-proxy:2
          volumeMounts: [{ name: signal, mountPath: /signal }]
          command: ["/bin/sh", "-c"]
          args:
            - |
              /cloud-sql-proxy --port=5432 proj:eu:inst &
              PID=$!
              until [ -f /signal/done ]; do sleep 1; done
              kill $PID; exit 0
      # and still no ordering guarantee: export may start before the proxy listens

After:

apiVersion: batch/v1
kind: Job
metadata: { name: nightly-export }
spec:
  template:
    spec:
      restartPolicy: Never
      initContainers:
        - name: sql-proxy
          image: gcr.io/cloud-sql-connectors/cloud-sql-proxy:2
          restartPolicy: Always            # <-- native sidecar
          args: ["--port=5432", "proj:eu:inst"]
          startupProbe:
            tcpSocket: { port: 5432 }
            periodSeconds: 1
      containers:
        - name: export
          image: myapp/export:2.1
          args: ["--db=127.0.0.1:5432"]

No shared volume, no shell wrapper, no kill, and the export container is guaranteed not to start until the proxy's startup probe passes.

Why It Was Added

Three recurring classes of bug, all caused by the fact that a regular container is just "a container" with no declared role:

  1. Jobs that never complete. A service-mesh proxy, log shipper or DB proxy has no exit condition. Pod stays Running, Job stays active, concurrencyPolicy: Forbid blocks the next schedule, and ttlSecondsAfterFinished never triggers so pods accumulate.
  2. Startup races. Regular containers start in parallel. Your app's first request goes out before the mesh proxy has its config, producing a burst of connection-refused errors at every rollout. The workaround was retry loops or postStart hooks that sleep.
  3. Shutdown races. On termination, all regular containers get SIGTERM at once. The log shipper dies while the app is still writing its last lines, so the logs explaining a crash are exactly the ones you lose.

Native sidecars encode the lifecycle relationship the mesh/logging ecosystem had been faking with injected shell scripts for years.

How It Works Underneath

The kubelet's container lifecycle logic gained a distinct class. A sidecar is stored in initContainers (so it inherits ordered startup) but with restartPolicy: Always overriding the pod-level policy (so it isn't expected to exit).

Sequence for a pod with one sidecar and one regular container:

sequenceDiagram
    participant K as kubelet
    participant S as sidecar (initContainer, restartPolicy Always)
    participant A as app container (regular)
    participant API as API server
    K->>S: start
    S-->>K: startupProbe OK -> "Started"
    K->>A: start (blocked until sidecar Started)
    A-->>K: exit 0
    Note over K: all regular containers terminated
    K->>S: SIGTERM (reverse declaration order)
    S-->>K: exited
    K->>API: Pod phase = Succeeded
    API-->>API: Job .status.succeeded += 1

Key mechanics that follow from this:

Should You Adopt It Yet

Yes, if you are on 1.29+ — and unreservedly on 1.33+ where it is GA. This is the intended mechanism now; Istio (ambient/sidecar injection), Linkerd, and the major log shippers already support emitting native sidecars.

Costs and caveats:

Migration Notes

Incremental path — migrate Jobs and CronJobs first, since they're where the payoff is largest:

  1. Find the victims. Any Job pod that's been Running far longer than its work should take:
kubectl get pods -A --field-selector=status.phase=Running \
  -o json | jq -r '.items[] | select(.metadata.ownerReferences[]?.kind=="Job") |
  "\(.metadata.namespace)/\(.metadata.name) \(.status.containerStatuses|map(.name+"="+(.state|keys[0]))|join(","))"'

A line with one terminated and one running is the exact bug.

  1. Grep for the hacks you can now delete: until [ -f, /signal/done, shareProcessNamespace: true, pkill, quitquitquit (the Istio/cloud-sql-proxy shutdown endpoints), and lifecycle.postStart sleeps used to fake ordering.

  2. Move the container definition verbatim from containers to initContainers, add restartPolicy: Always, and add a startupProbe if downstream containers depend on it being ready — ordering is guaranteed on Started, which without a probe means "process launched", not "port listening".

  3. Re-check resources. Because sidecar requests now sum with app requests, previously-fitting pods may go Pending. Compare kubectl describe node allocatable against the new totals before rolling cluster-wide.

  4. Update guardrails: policy rules, PodSpec validators, and dashboards that enumerate spec.containers must also walk spec.initContainers.

What breaks loudly: nothing at apply time on 1.29+. What breaks quietly: policies and log pipelines that no longer see the sidecar. Check those before you celebrate the green Job.

Key takeaway: Declare sidecars as initContainers with restartPolicy: Always — the kubelet then starts them first, keeps them running, and terminates them automatically once your app containers exit, so Jobs actually complete.

Real-world challenge

A nightly Kubernetes CronJob has been silently failing to complete for two weeks. `kubectl get pods` shows the pod as `Running` with `1/2` containers ready — the migration container shows `Terminated (Exit Code: 0)` in describe output, but a `cloud-sql-proxy` container is still Running. Because the Job never reaches Complete, `concurrencyPolicy: Forbid` blocks the next run, and you now have 14 days of missed migrations. What's happening and how do you fix it?

Diagnosis

A Job completes when all containers in the pod terminate. The proxy is a long-running server with no exit condition, so the pod never finishes and Forbid starves every subsequent schedule.

Confirm with:

kubectl get pod $POD -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.state}{"\n"}{end}'

You'll see migrate: terminated(0) next to cloud-sql-proxy: running.

The old hacks — a shared emptyDir sentinel file plus until [ -f /done ]; do sleep 1; done; exit 0 in the proxy, or making the app container kill PID 1 of the sidecar via shared process namespace. Both are fragile and mask real failures.

The fix (cluster on 1.29+, GA in 1.33): move the proxy into initContainers with restartPolicy: Always.

spec:
  template:
    spec:
      restartPolicy: Never
      initContainers:
        - name: cloud-sql-proxy
          image: gcr.io/cloud-sql-connectors/cloud-sql-proxy:2
          restartPolicy: Always      # makes it a sidecar
          args: ["--port=5432", "proj:region:inst"]
      containers:
        - name: migrate
          image: myapp/migrate:1.4

The kubelet starts the proxy first, waits for it to be started, runs the migration, then SIGTERMs the proxy once migration exits. Job goes Complete.

Cleanup: delete the stuck pods so Forbid releases, and add ttlSecondsAfterFinished plus an alert on kube_job_status_active staying > 0 longer than expected runtime.