Kubernetes CrashLoopBackOff: how to read it, find the cause, and fix it
CrashLoopBackOff is a symptom, not a cause. How the restart backoff works, what exit codes tell you, the kubectl ladder for finding the real failure, and how to apply a bounded fix you actually reviewed.
CrashLoopBackOff is the most recognisable failure state in Kubernetes and one of the most misread. It is not an error code from your application. It is Kubernetes telling you that a container keeps exiting and that the kubelet is now waiting longer between restart attempts. The cause is somewhere else entirely — usually in the logs of the instance that already died.
This guide is the operator ladder: what the state actually means, how to read the exit code, how to get the logs of the crashed container rather than the starting one, and how to make a bounded change you reviewed first.
What CrashLoopBackOff actually means
With the default restartPolicy of Always, the kubelet restarts a container that exits — for any reason, including a clean exit 0. If it keeps exiting, the kubelet applies an exponential backoff between attempts instead of hammering the node. The Pod reports CrashLoopBackOff while it waits.
- The container started — this is not a scheduling or image problem (that would be Pending or ImagePullBackOff)
- The restart count climbs and the gap between restarts grows
- Kubernetes is behaving correctly; your container is the thing that is unhappy
- A container that exits 0 immediately still loops, because Always means always
Read the exit code first
The exit code narrows the search dramatically before you read a single log line. It lives under the container's Last State in describe output.
| Exit code | Usually means | Where to look next |
|---|---|---|
| 1 | Application error — unhandled exception, failed startup check | Logs of the previous container |
| 137 | SIGKILL — most often OOMKilled at the memory limit | Last State reason and memory limits |
| 143 | SIGTERM — terminated, often during shutdown handling | Probe config and graceful shutdown code |
| 126 / 127 | Command not executable or not found | Image entrypoint, command, and PATH |
| 0 | Clean exit, but restartPolicy keeps restarting it | Whether this should be a Job instead of a Deployment |
If you see 137, you are probably not debugging a crash loop at all — you are debugging a memory kill that presents as one. That has its own ladder.
The kubectl ladder
The single most common mistake is running kubectl logs and seeing nothing useful, because that returns the currently starting container — not the one that crashed. You almost always want --previous.
Confirm the state, then read the right logs
kubectl get pods -n payments
kubectl describe pod -l app=api -n payments
# Containers → Last State:
# Reason: Error
# Exit Code: 1
# logs of the instance that actually died
kubectl logs -l app=api -n payments --previous --tail=100- Scope — which Deployment, namespace, and kubeconfig context
- Status — restart count, Ready, Last State reason and exit code
- Logs — always with --previous for a looping container
- Events — Back-off restarting failed container, probe failures, image errors
- Config — env vars, mounted Secrets and ConfigMaps, entrypoint
- Dependencies — is it dying because something it connects to is unreachable?
Events tell you what the kubelet is doing
kubectl get events -n payments --sort-by='.lastTimestamp' | tail -20
# Warning BackOff Back-off restarting failed container apiThe five causes that cover most crash loops
1. A dependency is not reachable
The container starts, tries to connect to a database or cache, fails, and exits. The log line is usually explicit — connection refused, no such host, timeout. Check that the Service name resolves and that the dependency is actually Ready before you touch the crashing workload.
2. Missing or wrong configuration
A required environment variable is unset, a Secret key was renamed, or a mounted ConfigMap does not have the file the app expects. These fail fast on startup, which is why the logs are short and the restart count is high.
3. Memory limit too low
Exit code 137 with reason OOMKilled. Raising the limit may be correct, or it may be hiding a leak. Compare observed usage to the limit before doubling anything.
4. A bad probe configuration
A liveness probe that starts checking before the app can answer will kill a healthy-but-slow container forever. If the app works when you exec into it but the kubelet keeps restarting it, suspect initialDelaySeconds, the probe path, or the port.
5. A bad release
It worked an hour ago. Nothing about the cluster changed. In that case the fastest safe action is usually to roll back, then debug the image on your own time rather than during the incident.
Reproduce it on purpose before you need to
The best time to practise this ladder is when nothing is actually on fire. kprompt-examples ships a CrashLoopBackOff fixture: the api container logs a connection attempt, reports connection refused to its database, then exits 1 — exactly the shape you meet in production.
A real crash loop in a kind cluster
git clone https://github.com/kprompt/kprompt-examples.git
cd kprompt-examples
make up && make break SCENARIO=01-crashloop && make verify
kubectl describe pod -l app=api -n payments
kubectl logs -l app=api -n payments --previous
make fix SCENARIO=01-crashloopAsking in natural language instead
The ladder above is mechanical, which is exactly why it is worth compiling. kprompt's explain path walks Deployment → ReplicaSet → Pods → Events → Logs and reports what it found, including the previous container's output. Reads run immediately — there is nothing to approve when nothing changes.
Explain is read-only
kprompt "explain why api is crashing" -n payments
kprompt "logs api" -n payments
kprompt "why is api not ready" -n paymentsWhen the right fix is a mutation — roll back, raise a limit, correct replicas — it becomes a plan with a diff and a risk verdict that you approve. That boundary is the point: the model can be wrong about the cause, and you still see exactly what would change before it changes.
A rollback you review first
$ kprompt "rollback api" -n payments
Plan
1. rollout undo Deployment/api
Risk: medium
Apply? [y/N]What not to do
- Do not delete the Pod and hope — the Deployment recreates it and the loop returns
- Do not remove the liveness probe to stop the restarts; you are deleting the alarm, not the fire
- Do not raise memory limits reflexively unless the exit code and reason actually say OOMKilled
- Do not read kubectl logs without --previous and conclude there are no logs
- Do not approve an AI-suggested plan you have not sanity-checked against the evidence
Related reading
For memory kills specifically, see the OOMKilled guide. For a wider prompt catalogue across failure modes, see the error prompt playbook. For always-on namespace watching that groups restarts into one Incident instead of paging per restart, see the Observe agent docs.
Related posts
Kubernetes OOMKilled: how to detect memory kills, raise limits, and avoid guesswork
A practical guide to OOMKilled in Kubernetes — exit 137, Last State, memory requests vs limits, kubectl checks, and how kprompt explain can suggest a reviewable memory patch.
Read articleHow to troubleshoot Kubernetes: deployments, pods, and crash loops from the terminal
A practical guide to Kubernetes troubleshooting — CrashLoopBackOff, deployments not ready, image pull errors, and rollbacks — using kubectl workflows and natural-language explains with kprompt.
Read articlekubectl vs K9s: when to use each (and why you keep both)
A head-to-head for operators: kubectl is the precise API client and scripting language, K9s is a live terminal UI over the same API. Which one to reach for during an incident, in CI, and while learning Kubernetes.
Read article