Skip to content
Hosting Operations9 min read

How to manage Kubernetes clusters in 2024?: Troubleshooting Checklist

Systematic troubleshooting guide for Kubernetes management platforms. Common failure symptoms, diagnostic checks, and concrete fixes with quick-reference

Written by Abdul AbrorTechnical Hosting Support Engineer
a computer generated image of a computer
On this page

TL;DR — Key takeaways

  • Most Kubernetes failures trace to resource exhaustion, misconfigured networking, or permission issues that can be diagnosed systematically using kubectl and native cluster monitoring.
  • A structured troubleshooting workflow starts with pod status checks, then moves through container logs, resource constraints, networking validation, and configuration verification before escalating to node-level diagnostics.
  • Kubernetes management platforms like Rancher, OpenShift, and Lens simplify cluster operations by providing centralized dashboards, but understanding core kubectl diagnostics remains essential for effective troubleshooting.
  • Always verify namespace context, RBAC permissions, and resource quotas before assuming application-level failures in multi-tenant Kubernetes environments.

Kubernetes management platforms have matured significantly, offering robust tools for cluster orchestration, monitoring, and lifecycle management. Yet even with sophisticated platforms like Rancher, OpenShift, or cloud-managed solutions, cluster administrators and support engineers face recurring troubleshooting scenarios that require systematic diagnosis.

This guide provides a practical troubleshooting framework for Kubernetes clusters, covering the most common failure symptoms and their resolution paths. Whether you're managing on-premises clusters or cloud-hosted infrastructure, these diagnostic steps apply across kubernetes management platforms and help you resolve issues efficiently without guessing.

Common Kubernetes Failure Symptoms

Before diving into diagnostics, recognize these failure patterns that account for the majority of Kubernetes cluster issues. Each symptom points toward specific diagnostic paths covered in later sections.

  • Pods stuck in Pending state: Indicates resource constraints, scheduling failures, or persistent volume binding issues
  • CrashLoopBackOff status: Application failures, misconfigured probes, or missing dependencies causing repeated restarts
  • ImagePullBackOff errors: Registry authentication problems, network connectivity issues, or incorrect image references
  • Service connectivity failures: DNS resolution problems, network policy restrictions, or service selector mismatches
  • Node NotReady status: Kubelet failures, resource exhaustion on the node, or network partition from the control plane
  • Resource quota exceeded messages: Namespace limits reached or resource requests exceeding available capacity
  • Slow or hanging kubectl commands: API server performance issues, etcd problems, or control plane resource constraints

Initial Diagnostic Checks: Pod and Container Level

For intermittent failures, use 'kubectl logs <pod-name> --since=1h' to focus on recent events. If pods are being evicted or terminated unexpectedly, check for OOMKilled status in the 'State' field of 'kubectl describe pod' output, which indicates memory limit breaches.

  • Check pod status: 'kubectl get pods -n <namespace>' shows running state, restarts, and age. Look for non-Running states or high restart counts
  • Describe the pod: 'kubectl describe pod <pod-name> -n <namespace>' displays events, scheduling decisions, and resource allocation. The Events section at the bottom often reveals the root cause
  • Inspect container logs: 'kubectl logs <pod-name> -n <namespace>' for single-container pods, or add '-c <container-name>' for multi-container pods. Use '--previous' flag to view logs from crashed containers
  • Check resource requests and limits: Verify that the pod's resource specifications match available node capacity using 'kubectl describe pod' output
  • Validate image availability: Confirm the image name and tag are correct, and test pull access from a node if ImagePullBackOff occurs

Network and Service Connectivity Diagnostics

If using istio or other service mesh implementations, remember that mesh proxy sidecars add another layer. Check sidecar injection status and envoy logs separately from application container logs. Network policies interact with service mesh rules, so both must align.

  • Verify service endpoints: 'kubectl get endpoints <service-name> -n <namespace>' confirms whether the service has discovered backing pods. Empty endpoints indicate selector mismatch or unhealthy pods
  • Test service ClusterIP directly: From a debug pod, use 'curl http://<cluster-ip>:<port>' to bypass DNS and test service routing
  • Check network policies: 'kubectl get networkpolicy -n <namespace>' lists active policies. Overly restrictive policies often block legitimate traffic in multi-tenant environments
  • Validate ingress configuration: For external access issues, verify ingress controller logs and confirm hostname/path mappings match service definitions
  • Review CNI plugin health: Check that the container network interface plugin (Calico, Cilium, Flannel, etc.) is running on all nodes

Node-Level and Cluster Resource Analysis

For production clusters, establish baseline metrics before troubleshooting so you can identify deviations. Most kubernetes management platforms include built-in monitoring dashboards for node metrics, but understanding the underlying kubectl commands ensures you can diagnose issues even when dashboards are unavailable.

  • Review resource utilization: 'kubectl top nodes' shows current CPU and memory usage. Compare against capacity to identify overcommitted nodes
  • Check pod distribution: 'kubectl get pods -o wide --all-namespaces' reveals whether pods are balanced or concentrated on specific nodes
  • Verify kubelet operation: SSH to problematic nodes and check 'systemctl status kubelet' and 'journalctl -u kubelet -n 100' for errors
  • Inspect disk usage: Nodes with full disks prevent new pods from scheduling. Check '/var/lib/docker' or '/var/lib/containerd' and pod log volumes
  • Review namespace resource quotas: 'kubectl get resourcequota -n <namespace>' and 'kubectl describe resourcequota' show quota limits and current usage

Configuration and RBAC Verification

When working with multiple kubernetes management platforms or migrating between them, configuration drift is common. Use 'kubectl diff -f <manifest.yaml>' before applying changes to preview the impact and catch configuration errors before they affect running workloads.

  • Validate RBAC bindings: Check that Roles, ClusterRoles, RoleBindings, and ClusterRoleBindings grant necessary permissions to service accounts used by your workloads
  • Test API server access: From within a pod, use 'curl -k https://kubernetes.default.svc/api/v1/namespaces' with the mounted service account token to confirm API connectivity
  • Review security contexts: Overly restrictive pod security policies or security contexts can prevent containers from functioning. Check 'securityContext' in pod definitions
  • Verify ConfigMaps and Secrets: Confirm referenced ConfigMaps and Secrets exist in the correct namespace and contain expected keys using 'kubectl get configmap' and 'kubectl get secret'
  • Check admission webhooks: Failed validation or mutation webhooks silently block resource creation. Review admission controller logs in the kube-system namespace

Control Plane and Storage Subsystem Checks

For managed Kubernetes services (EKS, GKE, AKS), control plane access is limited. Focus on persistent volume claims and storage class configuration instead. Contact cloud provider support if you suspect control plane issues beyond your diagnostic scope. Always maintain etcd backups before performing cluster maintenance—data loss from etcd corruption is not recoverable without backups.

  • Review API server logs: 'kubectl logs -n kube-system kube-apiserver-<node-name>' reveals authentication failures, rate limiting, and performance issues
  • Check etcd health: For clusters with direct etcd access, use 'etcdctl endpoint health' and 'etcdctl alarm list' to verify data store integrity
  • Validate scheduler operation: 'kubectl logs -n kube-system kube-scheduler-<node-name>' shows why pods aren't being scheduled to nodes
  • Inspect persistent volume status: 'kubectl get pv' and 'kubectl get pvc -A' identify storage binding failures. Pending PVCs block pod scheduling
  • Test storage class provisioning: Create a test PVC to verify that dynamic provisioning works for your storage backend

Using Kubernetes Management Platforms for Troubleshooting

Despite these advantages, master the underlying kubectl diagnostic commands. Management platform availability depends on network connectivity and authentication systems that may fail during cluster incidents. Command-line proficiency ensures you can troubleshoot under any conditions, and many support escalations require kubectl output rather than screenshots.

  • Centralized log aggregation: View logs from multiple pods and containers without repeated kubectl commands. Search across namespaces to identify patterns
  • Resource visualization: Graphical displays of CPU, memory, and disk usage help identify resource contention more quickly than tabular kubectl output
  • Built-in terminal access: Execute diagnostic commands directly from the management interface without separate SSH or kubeconfig setup
  • Configuration comparison: Some platforms highlight configuration drift between environments or show changes over time to identify recent modifications that caused failures
  • Alert integration: Management platforms often include alerting rules that proactively notify you of common failure conditions before they escalate

Quick troubleshooting checklist

  • Confirm you're working in the correct namespace context before running diagnostic commands
  • Check pod status and recent events using 'kubectl get pods' and 'kubectl describe pod'
  • Review container logs with 'kubectl logs', including '--previous' flag for crashed containers
  • Verify service endpoints match backing pods using 'kubectl get endpoints'
  • Test DNS resolution from within the cluster using a debug pod with network tools
  • Check node health and resource utilization with 'kubectl get nodes' and 'kubectl top nodes'
  • Review namespace resource quotas and limits that may block pod scheduling
  • Validate RBAC permissions for service accounts used by failing workloads
  • Confirm ConfigMaps and Secrets exist in the correct namespace with expected keys
  • Check persistent volume claim status and storage class provisioning configuration
  • Review control plane component logs in kube-system namespace for API server and scheduler errors
  • Test resource creation with 'kubectl diff' before applying changes to production clusters
  • Verify network policies and service mesh configurations aren't blocking legitimate traffic
  • Check kubelet status on nodes showing NotReady or pressure conditions
  • Maintain current etcd backups before performing any cluster maintenance operations

FAQ

What does it mean when a Kubernetes pod is stuck in Pending status?

A pod in Pending status indicates it hasn't been scheduled to a node yet. Common causes include insufficient CPU or memory resources on available nodes, unsatisfied node selectors or affinity rules, pending persistent volume claims that can't be bound, or resource quota limits reached in the namespace. Use 'kubectl describe pod' to see the specific scheduling failure reason in the Events section.

How do I troubleshoot CrashLoopBackOff errors in Kubernetes?

CrashLoopBackOff means the container is crashing repeatedly after starting. Check container logs using 'kubectl logs <pod-name> --previous' to see output from the crashed instance. Common causes include application startup failures due to missing environment variables or configuration, failed liveness or readiness probes with incorrect settings, insufficient memory causing OOMKilled terminations, or missing dependencies like database connectivity. Fix the underlying application or configuration issue, then redeploy.

Why can't my Kubernetes pods reach services in other namespaces?

Pods access services in other namespaces using the fully qualified DNS name format: '<service-name>.<namespace>.svc.cluster.local'. Verify DNS resolution works by testing from a debug pod, then check that network policies in both source and destination namespaces allow the traffic. Review service selectors to confirm they match the backing pod labels, and verify the service has endpoints using 'kubectl get endpoints'. RBAC policies don't affect network connectivity between pods, but service mesh configurations or CNI plugin rules might.

What are the essential kubectl commands for Kubernetes troubleshooting?

The core troubleshooting commands are: 'kubectl get pods -n <namespace>' for status overview, 'kubectl describe pod <name>' for detailed events and configuration, 'kubectl logs <pod-name>' for container output, 'kubectl get events -n <namespace> --sort-by=.lastTimestamp' for recent cluster events, 'kubectl top nodes' and 'kubectl top pods' for resource usage, and 'kubectl get all -n <namespace>' for all resources in a namespace. These commands diagnose 90% of common cluster issues when used systematically.

How do kubernetes management platforms improve cluster troubleshooting compared to kubectl alone?

Kubernetes management platforms like Rancher, OpenShift, and Lens provide centralized dashboards that aggregate cluster state, resource metrics, and logs from multiple sources without repeated command execution. They offer graphical resource usage visualization, built-in terminal access, configuration comparison tools, and alert integration that makes pattern recognition faster than command-line analysis. However, kubectl remains essential for detailed diagnostics, automation, and troubleshooting when management platform interfaces are unavailable during incidents.