Kubernetes Cluster Troubleshooting: 8 CrashLoop Fixes
Fix CrashLoopBackOff, ImagePullBackOff, and OOMKilled errors in your Kubernetes cluster. Eight proven steps with kubectl commands and security checks.

On this page
- Understanding the Kubernetes crash loop threat model
- Diagnostic audit: what to check before changing anything
- Fix 1: Resolve ImagePullBackOff with correct registry secrets
- Fix 2: Correct application errors causing immediate exit
- Fix 3: Set resource limits to prevent OOMKilled errors
- Fix 4: Harden security contexts without breaking workloads
- Fix 5: Implement Pod Security Standards at the namespace level
- Fix 6: Audit RBAC permissions for pod service accounts
- Fix 7: Configure liveness and readiness probes correctly
- Fix 8: Set resource quotas and limit ranges per namespace
- Verifying your security configuration is actually working
TL;DR — Key takeaways
- CrashLoopBackOff indicates your container exits immediately after starting—check application logs first with kubectl logs to find the actual error
- ImagePullBackOff means the cluster cannot pull your container image; verify registry credentials exist as Secrets and image names match exactly
- Resource limits prevent OOMKilled errors but introduce new failure modes—set requests below limits and monitor actual usage before applying cluster-wide
- Security contexts that block root access or enforce read-only filesystems will crash legacy applications—test one pod before rolling out restrictions
A pod stuck in CrashLoopBackOff is one of the first things I investigate when a hosting customer reports their application is down. The error name sounds worse than it is—it just means the container exited and Kubernetes keeps trying to restart it. The real problem is in your application logs or configuration.
In this guide I'll walk through eight fixes for the most common cluster errors, starting with diagnostic commands you run first and ending with security hardening that prevents these issues from recurring. Each fix includes the kubectl command, what to look for in the output, and the YAML changes needed.
Understanding the Kubernetes crash loop threat model
When a pod crashes repeatedly, you face two separate problems. The immediate operational issue is downtime—your app isn't serving traffic. The security problem is less obvious: crash loops often expose configuration mistakes that attackers can exploit.
Pods that crash because of missing credentials might log sensitive connection strings. Containers running as root can write to the host filesystem if you misconfigured volume mounts. OOMKilled pods sometimes dump memory contents to logs before dying.
The threat model here is information disclosure through error messages and privilege escalation through loose security contexts. Both happen most often during the initial deployment when teams rush to 'just make it work' and defer hardening.
Diagnostic audit: what to check before changing anything
Start kubernetes cluster troubleshooting by collecting facts, not guessing. Run kubectl get pods -n <namespace> and note the pod status. CrashLoopBackOff, ImagePullBackOff, and Error are the states you're hunting.
Next run kubectl describe pod <pod-name> -n <namespace>. Scroll to the Events section at the bottom—it shows exactly what failed. You'll see messages like 'Back-off pulling image' or 'Container exited with code 1'.
For application-level errors, pull the logs: kubectl logs <pod-name> -n <namespace> --previous. The --previous flag shows logs from the crashed container before Kubernetes restarted it. Most crash causes appear in the last 20 lines.
- Check pod events first with kubectl describe—errors appear in chronological order
- Exit code 0 means clean shutdown, 1 means application error, 137 means OOMKilled
- If logs are empty, the container never started—likely an ImagePullBackOff issue
- Run kubectl get events --sort-by='.lastTimestamp' to see cluster-wide failures
Fix 1: Resolve ImagePullBackOff with correct registry secrets
Verify the secret exists with kubectl get secret regcred -n <namespace> -o yaml. If you see the secret but pulls still fail, check that the image name exactly matches the repository path—extra slashes or missing tags cause silent failures.
- imagePullSecrets:
- - name: regcred
Fix 2: Correct application errors causing immediate exit
When kubectl logs shows your application starting then crashing, the problem is in your code or its runtime dependencies. I've seen this most often with missing environment variables, unreachable databases, and file permission errors.
Check your ConfigMap and Secret references first. Run kubectl get configmap -n <namespace> and kubectl get secret -n <namespace> to confirm they exist. Then verify your deployment mounts them correctly:
Look for envFrom or env sections in your pod spec that reference configMapRef or secretRef. A typo in the name means your app gets empty variables and crashes during init.
- Test container locally with the same env vars: docker run --env-file .env <image>
- Add verbose logging to your app's startup sequence so crashes are obvious
- Use an init container to validate dependencies before starting the main container
- Set terminationMessagePolicy: FallbackToLogsOnError to capture crash details
Fix 3: Set resource limits to prevent OOMKilled errors
Requests determine scheduling—Kubernetes won't place your pod on a node without available capacity. Limits enforce hard caps. Setting requests too high wastes resources; setting limits too low causes crashes.
After applying new limits, monitor for 24 hours. Memory leaks show up as gradually increasing usage until the next OOMKill.
- resources:
- requests:
- memory: '512Mi'
- cpu: '250m'
- limits:
- memory: '1Gi'
- cpu: '500m'
Fix 4: Harden security contexts without breaking workloads
Read-only root breaks apps that write temp files to /tmp. The fix is mounting an emptyDir volume at /tmp so the app can still write there but not to system directories.
- allowPrivilegeEscalation: false
- readOnlyRootFilesystem: true
- capabilities:
- drop:
- - ALL
Fix 5: Implement Pod Security Standards at the namespace level
Pod Security Standards replaced PodSecurityPolicies in Kubernetes 1.25. They define three levels: privileged (unrestricted), baseline (prevents known escalations), and restricted (deeply hardened).
Apply baseline to existing namespaces and restricted to new ones. Label the namespace:
kubectl label namespace <namespace> pod-security.kubernetes.io/enforce=baseline pod-security.kubernetes.io/audit=restricted pod-security.kubernetes.io/warn=restricted
Enforce mode blocks non-compliant pods. Audit mode logs violations without blocking. Warn mode shows warnings in kubectl output. Use warn first to discover what will break.
- Baseline blocks hostPath volumes, hostNetwork, and privileged containers
- Restricted additionally requires runAsNonRoot and blocks all capabilities
- Check violations with kubectl get events --field-selector reason=FailedCreate
- Exempt specific pods by adding a label: pod-security.kubernetes.io/enforce=privileged
Fix 6: Audit RBAC permissions for pod service accounts
Never grant cluster-admin or edit to application ServiceAccounts. Scope permissions to exactly what the app needs—usually just 'get' and 'list' on specific resource types.
- kubectl create serviceaccount myapp-sa -n <namespace>
- Create a Role with rules for configmaps, secrets, or other resources
- Bind the Role to the ServiceAccount with a RoleBinding
- Reference the ServiceAccount in your deployment: serviceAccountName: myapp-sa
Fix 7: Configure liveness and readiness probes correctly
Set initialDelaySeconds longer than your app's startup time. If startup takes 20 seconds and you set 5, the probe will fail before the app is ready and Kubernetes will restart it immediately.
- livenessProbe:
- httpGet:
- path: /healthz
- port: 8080
- initialDelaySeconds: 30
- periodSeconds: 10
- readinessProbe:
- httpGet:
- path: /ready
- port: 8080
- initialDelaySeconds: 10
- periodSeconds: 5
Fix 8: Set resource quotas and limit ranges per namespace
Any pod created without explicit resource definitions inherits these defaults. That prevents unbounded resource usage from lazy deployment specs.
- apiVersion: v1
- kind: LimitRange
- metadata:
- name: default-limits
- namespace: <namespace>
- spec:
- limits:
- - default:
- memory: 512Mi
- cpu: 500m
- defaultRequest:
- memory: 256Mi
- cpu: 250m
- type: Container
Verifying your security configuration is actually working
After applying fixes, test that they work. Deploy a known-bad pod spec to confirm it gets blocked. Create a test deployment that violates your Pod Security Standard—it should fail with a clear error message.
Run kubectl auth can-i <verb> <resource> --as=system:serviceaccount:<namespace>:<sa-name> to test RBAC permissions without actually running a pod. For example, kubectl auth can-i get secrets --as=system:serviceaccount:default:myapp-sa should return 'yes' only if you granted that permission.
Check that resource limits prevent overconsumption by running a stress test container. Deploy a pod with a memory limit of 100Mi, then run a process inside that allocates 200Mi—it should get killed with exit code 137.
For read-only filesystems, exec into a hardened pod and try to write to /etc or /usr. It should fail with 'Read-only file system'. If it succeeds, your security context isn't applied correctly.
- Use kubectl describe pod to verify securityContext fields appear in the running pod spec
- Check that ServiceAccount tokens are mounted only when explicitly requested
- Scan images with trivy or similar tools to catch vulnerabilities before deployment
- Run kube-bench to audit cluster-level security against CIS benchmarks
Quick troubleshooting checklist
- Run kubectl describe pod <name> and check the Events section for error messages
- Verify image pull secrets exist in the target namespace with kubectl get secrets
- Check container exit codes in pod status (exit code 137 means OOMKilled, 1 means application error)
- Review resource requests and limits in your deployment YAML—ensure requests are realistic
- Test security contexts (runAsNonRoot, readOnlyRootFilesystem) on a single replica before scaling
- Audit RBAC permissions if pods cannot access ConfigMaps or Secrets they need
- Enable Pod Security Standards at the namespace level and validate existing workloads against them
- Set up resource quotas per namespace to prevent one team from starving cluster capacity
FAQ
What does CrashLoopBackOff mean in Kubernetes?
CrashLoopBackOff means your container started successfully but exited immediately, and Kubernetes is waiting longer between each restart attempt. The actual cause is in the application logs—run kubectl logs <pod-name> to see why the process died. Common causes include missing environment variables, failed database connections, or permission errors when the app tries to write to a read-only filesystem.
How do I fix ImagePullBackOff errors?
ImagePullBackOff happens when Kubernetes cannot pull the container image from your registry. First verify the image name and tag are correct in your deployment spec—typos are common. If the image is in a private registry, create an imagePullSecrets reference in your pod spec that points to a Secret containing registry credentials. Run kubectl describe pod to see the exact error message from the container runtime.
Why does my pod keep getting OOMKilled?
OOMKilled (exit code 137) means your container used more memory than its limit and the kernel killed it. Check current limits with kubectl describe pod and compare against actual usage shown in kubectl top pod. Increase the memory limit if usage is legitimate, or fix memory leaks in your application. Setting limits too high wastes cluster resources, so monitor real usage over several hours before deciding.
Related articles
- Hosting OperationsSelf-Hosted App Deployment Fails? Check DNS, SSL, Reverse Proxy, and Logs FirstTroubleshoot failed self-hosted app deployments by checking DNS, SSL, reverse proxy routing, container status, logs, and ports.
- Hosting OperationsSelf-Hosted PaaS on a VPS: What to Check Before Installing Coolify, Dokploy, or CapRoverA hosting support checklist for preparing a VPS before installing self-hosted PaaS tools like Coolify, Dokploy, or CapRover.
- Hosting OperationsLinux Server Security Lessons from the Arch Linux Malware Package IncidentPractical Linux server security checklist for VPS admins after package malware concerns, with safe checks, rollback steps, and support guidance.
- Hosting OperationsAWS Lightsail Hong Kong VPS Latency: Practical Hosting Guide for IndonesiaLearn how to test AWS Lightsail Hong Kong VPS latency, compare regions, migrate safely, and troubleshoot hosting performance.
- Hosting OperationsCloudflare Tomorrow Watchlist: A Practical Hosting Operations GuidePractical Cloudflare troubleshooting checklist for DNS, SSL, caching, WAF, origin health, safe testing, and rollback planning.
- Hosting OperationsNetwork Safety Checklist for AI Agent Skills in Hosting OperationsAudit AI agent skills safely with network checks, secret protection, sandbox testing, rollback steps, and hosting support troubleshooting guidance.