Skip to content
Hosting Operations8 min read

Kubernetes Cluster Setup: 5 Configs That Prevent Downtime

Stop cluster outages before they start. Configure resource limits, health checks, and pod disruption budgets during kubernetes cluster setup.

Written by Abdul AbrorTechnical Hosting Support Engineer
a purple background with a black and blue circle surrounded by blue and green cubes
On this page

TL;DR — Key takeaways

  • Set memory and CPU limits on every pod to prevent resource exhaustion from crashing nodes
  • Configure liveness and readiness probes so Kubernetes stops routing traffic to unhealthy pods before users see errors
  • Use PodDisruptionBudgets to guarantee minimum replica counts during node drains and cluster upgrades
  • Deploy at least three replicas across multiple availability zones for workloads that cannot tolerate downtime
  • Enable audit logging and metrics collection from day one—you need historical data to diagnose incidents after they happen

Most cluster outages I've seen in production started during setup, not six months later when traffic spiked. A deployment without resource limits triggers an OOMKill cascade. A missing readiness probe routes requests to pods still booting. Node maintenance evicts every replica because no one set a disruption budget.

The fix isn't more monitoring or a bigger cluster. It's five configuration decisions you make before the first real workload goes live. This guide walks through each one with the exact YAML you need and the failure mode it prevents.

Why Resource Requests and Limits Matter

Every container in a Kubernetes pod should declare two values: a resource request and a resource limit. The request reserves capacity on the node. The limit caps consumption.

Without requests, the scheduler has no idea how much memory or CPU your pod needs. It might land ten pods on one node and leave others idle. Without limits, a memory leak or runaway process can starve every other workload on that node.

Here's a deployment snippet that sets both:

  • requests.memory: 256Mi — reserves 256 megabytes; the scheduler won't place this pod on a node with less free memory
  • limits.memory: 512Mi — kills the container if it tries to allocate more than 512 megabytes
  • requests.cpu: 100m — reserves 0.1 CPU cores
  • limits.cpu: 500m — throttles the container if it exceeds 0.5 cores

Liveness and Readiness Probes Stop Traffic to Broken Pods

A liveness probe tells Kubernetes when to restart a container. A readiness probe tells it when to add the pod to the service endpoint list.

In support tickets I handled, the usual culprit was a Java app that took 45 seconds to start but had initialDelaySeconds set to 10. Kubernetes marked the pod ready, the load balancer sent requests, users saw 502s.

Configure readiness to check actual application health, not just that the process is running. If your app depends on a database connection, the readiness endpoint should fail when that connection is down. Liveness can be simpler—just verify the process can still respond.

Example probe configuration:

  • readinessProbe with httpGet on /health, initialDelaySeconds 30, periodSeconds 5
  • livenessProbe with httpGet on /ping, initialDelaySeconds 60, periodSeconds 10, failureThreshold 3
  • Set initialDelaySeconds longer than your worst-case startup time; measure it with time curl in a test pod

PodDisruptionBudgets Protect Availability During Maintenance

A PodDisruptionBudget (PDB) sets the minimum number of replicas that must stay running during voluntary disruptions. Voluntary means node drains, cluster upgrades, or any operation triggered by kubectl drain.

Without a PDB, draining a node evicts all pods immediately. If all your replicas happen to be on that node, your service goes down. I've seen this exact scenario take out a payment API during a routine Kubernetes version upgrade.

Create a PDB for every workload that can't tolerate zero replicas. Here's the YAML:

  • minAvailable: 2 — at least two pods must stay running during the drain
  • Alternative: maxUnavailable: 1 — at most one pod can be evicted at a time
  • Match the selector to your deployment's labels
  • Test with kubectl drain --dry-run before applying in production

Spread Pods Across Nodes and Zones

By default, the scheduler might place all replicas on the same node or in the same availability zone. That's fine until the node reboots or the zone has a network partition.

Use pod anti-affinity rules to force distribution. The podAntiAffinity block in your deployment spec tells Kubernetes to avoid scheduling multiple replicas on the same node or in the same zone.

Here's the pattern:

  • topologyKey: kubernetes.io/hostname — spreads replicas across different nodes
  • topologyKey: topology.kubernetes.io/zone — spreads replicas across different availability zones
  • Combine both for maximum fault tolerance
  • Requires at least as many nodes or zones as your replica count, or pods will stay pending

Why You Need Logging and Metrics Before the First Incident

When a pod crashes at 3 AM, you need to know what happened. Kubernetes events expire after an hour. Container logs vanish when the pod is deleted. If you didn't ship logs and metrics to a centralized system, you're troubleshooting blind.

Enable kubelet metrics and deploy a logging agent (Fluentd, Fluent Bit, or Promtail) on every node. Point them at a storage system outside the cluster—if the cluster goes down, you still have the data.

Essential metrics to collect:

  • container_memory_working_set_bytes — actual memory usage, the value Kubernetes uses for eviction decisions
  • container_cpu_usage_seconds_total — CPU time consumed by each container
  • kube_pod_status_phase — tracks pod lifecycle states (Pending, Running, Failed)
  • node_memory_MemAvailable_bytes — free memory on each node
  • Ship application logs with structured JSON so you can filter by pod, namespace, and timestamp

How to Test Configuration Changes Safely

Every configuration change carries risk. A typo in a resource limit can make pods unschedulable. An incorrect probe endpoint can mark healthy pods as unready.

Test in a non-production namespace first. Use kubectl apply --dry-run=client to catch syntax errors. Use kubectl apply --dry-run=server to validate against the actual cluster API.

For changes that affect running workloads, use a rolling update strategy and watch pod status in real time with kubectl get pods -w. If new pods fail readiness checks, the rollout pauses automatically.

Always document the rollback command before applying the change. For a deployment, that's kubectl rollout undo deployment/your-app. For a PDB or config map, keep the previous YAML in version control and apply it if things go wrong.

  • kubectl diff -f manifest.yaml shows what will change before you apply it
  • kubectl apply --dry-run=server -f manifest.yaml validates against the cluster without making changes
  • kubectl rollout status deployment/your-app blocks until the rollout completes or fails
  • Set maxUnavailable: 1 in your deployment strategy so updates never take down more than one pod at a time

Quick Reference Table

Here's a summary of the five configurations, the failure mode each one prevents, and where to define it in your manifest.

  • Resource requests and limits — prevents OOMKills and node resource exhaustion — defined in spec.containers[].resources
  • Readiness and liveness probes — prevents traffic to unhealthy pods and restarts stuck containers — defined in spec.containers[].readinessProbe and livenessProbe
  • PodDisruptionBudget — prevents zero replicas during node drains and upgrades — separate PDB resource with selector matching deployment labels
  • Pod anti-affinity — prevents all replicas on the same node or zone — defined in spec.template.spec.affinity.podAntiAffinity
  • Centralized logging and metrics — provides historical data for post-incident analysis — deploy logging agent as DaemonSet and enable kubelet metrics endpoint

Common Mistakes to Avoid

Setting limits too low causes constant OOMKills even under normal load. Measure actual memory usage in staging with kubectl top pods and add 20% headroom.

Copying probe settings from examples without adjusting initialDelaySeconds for your app's real startup time will cause endless restart loops or premature ready signals.

Forgetting to create a PDB for stateful workloads like databases or message queues. A single evicted pod can lose data if it's the only replica with recent writes.

Using requiredDuringSchedulingIgnoredDuringExecution for anti-affinity when you don't have enough nodes. Pods will stay pending forever. Use preferredDuringSchedulingIgnoredDuringExecution instead to make it a soft constraint.

Quick troubleshooting checklist

  • Set requests and limits for CPU and memory in every deployment manifest
  • Add readinessProbe and livenessProbe to all application containers
  • Create PodDisruptionBudget resources for stateful and customer-facing workloads
  • Configure pod anti-affinity rules to spread replicas across nodes and zones
  • Enable kubelet metrics and ship logs to a centralized system outside the cluster
  • Test node drain and pod eviction before the first production deploy
  • Document rollback steps for configuration changes before applying them

FAQ

What happens if I don't set resource limits during Kubernetes cluster setup?

Without resource limits, a single pod can consume all available memory or CPU on a node, causing the kubelet to evict other pods or triggering an out-of-memory kill. The node may become unresponsive, and kube-apiserver will mark it NotReady. Set requests to reserve capacity and limits to prevent runaway consumption.

How do I configure health checks that actually prevent user-facing errors?

Use readinessProbe to pull pods out of service rotation before they fail and livenessProbe to restart containers stuck in a broken state. Point readiness checks at an endpoint that validates all dependencies—database connections, upstream APIs—not just HTTP 200 on the root path. Set initialDelaySeconds long enough for your app to finish startup.

What is a PodDisruptionBudget and when do I need one?

A PodDisruptionBudget (PDB) tells Kubernetes the minimum number of replicas that must stay running during voluntary disruptions like node drains or cluster upgrades. Without a PDB, kubectl drain can evict all replicas at once, causing downtime. Create a PDB for any workload that serves live traffic or holds stateful data.