Skip to content
Hosting Operations11 min read

Kubernetes Cluster Setup in 8 Steps + Performance Tuning

Set up a production Kubernetes cluster from bare metal and tune CPU limits, network CNI, and disk I/O to cut pod start times by 40%.

Written by Abdul AbrorTechnical Hosting Support Engineer
a purple background with a black and blue circle surrounded by blue and green cubes
On this page

TL;DR — Key takeaways

  • A bare-metal Kubernetes cluster requires eight core steps: hardware prep, OS hardening, container runtime install, control plane init, CNI deployment, worker node join, storage provisioner config, and ingress setup
  • CPU throttling from default cgroups v1 limits causes 60-80% of slow pod starts; switching to cgroups v2 and tuning cpu.max reduces startup time by 35-45%
  • Network plugin choice directly impacts latency—Cilium with eBPF cuts pod-to-pod overhead to under 0.2ms versus 1.5ms for standard Flannel VXLAN
  • Disk I/O contention on etcd's data directory kills API responsiveness; mount etcd on dedicated NVMe with noatime and monitor fsync latency below 10ms

Kubernetes cluster setup on bare metal still trips up teams who learned k8s through managed services. The control plane won't start if swap is on. Pods hang in ContainerCreating when you skip the CNI step. Join tokens expire in 24 hours.

I've walked dozens of support tickets through failed cluster bootstraps—usually a missing cgroup driver or a firewall rule blocking 6443. This guide covers the exact eight-step sequence to bring up a working cluster, then digs into the three performance bottlenecks that slow down every new deployment: CPU throttling, network overhead, and disk I/O contention on etcd. You'll see real tuning commands and the before/after metrics I collected on production hardware.

Step 1–3: Hardware Prep, Runtime Install, and OS Hardening

Start with three physical or VM nodes—one control plane, two workers. Each needs a unique hostname, MAC address, and product_uuid (check with cat /sys/class/dmi/id/product_uuid). Minimum 2 CPU cores and 2GB RAM per node. Kubernetes won't schedule on nodes with less than 1 CPU available after system overhead.

Disable swap immediately. Run swapoff -a then edit /etc/fstab to comment out any swap line. Kubelet refuses to start if swap is active—it breaks memory isolation guarantees.

Install a container runtime next. Containerd is the standard choice now that Docker shim is removed. On Ubuntu or Debian, install containerd.io from Docker's apt repo, then configure the systemd cgroup driver by editing /etc/containerd/config.toml. Set SystemdCgroup = true under [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runc.options]. Restart containerd after the edit.

Load required kernel modules: overlay and br_netfilter. Add them to /etc/modules-load.d/k8s.conf, then run modprobe overlay && modprobe br_netfilter. Set sysctl parameters for IP forwarding and bridge traffic: net.bridge.bridge-nf-call-iptables=1, net.ipv4.ip_forward=1, net.bridge.bridge-nf-call-ip6tables=1 in /etc/sysctl.d/k8s.conf. Apply with sysctl --system.

Step 4–5: Control Plane Init and CNI Deployment

Install kubeadm, kubelet, and kubectl from the Kubernetes apt or yum repository. Pin the version—mixing kubelet 1.28 with kubeadm 1.29 causes cryptic API errors.

Initialize the control plane with kubeadm init --pod-network-cidr=10.244.0.0/16 --apiserver-advertise-address=<NODE_IP>. The pod CIDR must match your CNI's expected range. Save the entire output—it contains the join command with a token valid for 24 hours.

Copy the admin kubeconfig to your home directory: mkdir -p $HOME/.kube && cp /etc/kubernetes/admin.conf $HOME/.kube/config && chown $(id -u):$(id -g) $HOME/.kube/config. Without this, kubectl commands fail with connection refused.

Deploy a CNI plugin within five minutes or pods will stay in Pending forever. For Flannel, apply the official manifest with kubectl apply -f https://github.com/flannel-io/flannel/releases/latest/download/kube-flannel.yml. Watch coredns pods with kubectl get pods -n kube-system until they reach Running. If they stay Pending, the CNI didn't install correctly—check kubectl describe pod and look for network plugin errors.

Step 6–8: Worker Join, Storage, and Ingress

On each worker node, run the kubeadm join command from the init output. It looks like kubeadm join <CONTROL_IP>:6443 --token <TOKEN> --discovery-token-ca-cert-hash sha256:<HASH>. If the token expired, generate a new one on the control plane with kubeadm token create --print-join-command.

Verify all nodes with kubectl get nodes. They should show Ready within 60 seconds. NotReady usually means the CNI isn't running on that node—check kubectl get pods -n kube-system -o wide and confirm a CNI pod is assigned to the worker.

Install a storage provisioner unless you only run stateless workloads. Longhorn is easy for small clusters—apply the manifest from their GitHub releases page. It creates a StorageClass named longhorn that provisions PersistentVolumes on local disks. For production, consider Rook-Ceph for distributed storage with replication.

Deploy an ingress controller last. Nginx-ingress is the default choice—install via Helm or the official manifests. Expose it as a NodePort service or configure MetalLB for LoadBalancer IPs on bare metal. Test by creating an Ingress resource pointing to a sample service and curling the ingress IP.

Performance Bottleneck 1: CPU Throttling from cgroups v1

Default cgroups v1 settings cause pods to throttle even when CPU is idle. The kernel enforces cpu.cfs_quota_us in 100ms periods, so a pod with 1 CPU limit can only use 100ms of CPU time per period. Bursty workloads hit this wall immediately—container startup, npm install, image builds.

Check throttling with kubectl top pods and compare CPU usage to limits. If usage stays pinned at exactly the limit but pods are slow, you're throttled. Also inspect /sys/fs/cgroup/cpu/kubepods/pod<UUID>/cpu.stat on the node—nr_throttled and throttled_time show how often the kernel blocked your containers.

Switch to cgroups v2 by adding systemd.unified_cgroup_hierarchy=1 to the kernel command line in /etc/default/grub, then run update-grub and reboot. Cgroups v2 uses cpu.max instead of cfs_quota and allows better burst allocation. After the reboot, configure kubelet and containerd to use systemd cgroup driver (already done if you followed step 3).

For high-priority pods, increase CPU limits or remove them entirely if you trust the workload. On a dedicated k8s cluster, I've seen pod start times drop from 12 seconds to 7 seconds just by switching to cgroups v2 and raising cpu.max on the kubelet cgroup itself to allow more scheduler overhead.

Performance Bottleneck 2: Network Plugin Overhead

Flannel VXLAN wraps every packet in UDP encapsulation, adding 50 bytes and 1-2ms latency. On a 1Gbps link that's acceptable. On 10GbE or 25GbE, you're leaving throughput on the table.

Measure baseline latency with iperf3 between two pods on different nodes. Start a server pod: kubectl run iperf-server --image=networkstatic/iperf3 --port=5201 -- -s. Exec into a client pod on another node and run iperf3 -c <SERVER_POD_IP> -t 30. Note the reported latency and throughput.

Cilium with eBPF skips iptables entirely and programs the kernel's XDP layer directly. Install Cilium via Helm with --set tunnel=disabled to use direct routing (requires L2 connectivity between nodes or BGP setup). After deployment, rerun the iperf3 test. On my hardware (Intel X550 NICs, kernel 6.1), latency dropped from 1.5ms to 0.18ms and throughput jumped from 7.2 Gbps to 9.4 Gbps.

Calico in IPIP or VXLAN mode adds similar overhead to Flannel. Switch Calico to native routing if your network allows it—set CALICO_IPV4POOL_IPIP to Never in the manifest. You'll need to handle routing between nodes yourself (BGP or static routes), but pod-to-pod packets travel as plain IP with no tunnel tax.

Performance Bottleneck 3: Disk I/O Contention on etcd

Etcd is the brain of your cluster. Every kubectl command, every pod schedule, every ConfigMap write hits etcd. It uses fsync on every commit to guarantee durability, so disk latency directly impacts API responsiveness.

Check etcd performance with etcdctl from inside the etcd pod (or from a control plane node if you installed etcd via kubeadm). Run etcdctl check perf to get write latency stats. Healthy etcd completes fsyncs in under 10ms. If you see 50ms or higher, your disk is the problem.

Mount etcd's data directory on dedicated NVMe storage. Edit the etcd static pod manifest at /etc/kubernetes/manifests/etcd.yaml on the control plane and change the hostPath volume to point to an NVMe mount like /mnt/etcd-nvme. Add the noatime mount option in /etc/fstab to skip access time updates—etcd doesn't need them and every write triggers an inode update otherwise.

If you can't dedicate a disk, at least separate etcd from the root filesystem. I've seen API request latency drop from 250ms to 40ms by moving etcd to a separate SSD. Also raise --quota-backend-bytes in the etcd command args (default is 2GB)—clusters with lots of ConfigMaps or Secrets hit this limit and start refusing writes. 8GB is a safe production value.

Monitoring and Ongoing Tuning

Deploy Prometheus and Grafana to track cluster metrics over time. Use kube-prometheus-stack from Helm—it includes prebuilt dashboards for CPU throttling, network throughput, and etcd latency.

Watch these metrics weekly. CPU throttling counters (container_cpu_cfs_throttled_seconds_total) tell you which namespaces need limit increases. Network transmit/receive error counters (node_network_transmit_errs_total) reveal NIC or CNI issues. Etcd fsync duration (etcd_disk_wal_fsync_duration_seconds) warns you before API slowdowns hit users.

Set resource requests on every production pod. Requests don't throttle but do reserve capacity—the scheduler won't pack pods so tightly that nodes run out of CPU. In support tickets, missing requests are the #1 cause of mystery evictions and OOMKills.

Tune node-level settings based on workload. Increase kernel.pid_max if you run lots of small containers. Raise fs.inotify.max_user_instances for log aggregators that watch many files. Test changes on a single worker node first, monitor for a day, then roll out clusterwide.

Rollback Boundaries and Safe Testing

Before making any kernel or cgroup change, snapshot your control plane node or back up /etc/kubernetes and /var/lib/etcd. A bad grub config or missing kernel module can brick the cluster.

Test performance changes on a single worker node by draining it first with kubectl drain <NODE> --ignore-daemonsets. Apply your tuning (cgroups v2, CNI swap, sysctl changes), then uncordon with kubectl uncordon <NODE> and compare pod performance to other workers. If metrics look good after 24 hours, roll out to the rest of the cluster.

Keep the previous kernel installed so you can boot back if cgroups v2 breaks something. On Ubuntu, install linux-image-generic-hwe to get newer kernels while keeping the old one as a fallback in grub.

Document every tuning parameter in a git repo alongside your kubeadm config. Three months from now, you won't remember why you set net.core.somaxconn to 4096, and the next engineer inheriting this cluster deserves to know.

Quick troubleshooting checklist

  • Verify each node meets 2 CPU / 2GB RAM minimum and has unique hostname, MAC, product_uuid
  • Disable swap with swapoff -a and comment swap line in /etc/fstab
  • Install containerd or CRI-O and configure systemd cgroup driver
  • Run kubeadm init on control plane node and save the join command output
  • Deploy a CNI plugin (Calico, Cilium, or Flannel) within 5 minutes of init
  • Join worker nodes using the kubeadm join token from control plane output
  • Install a CSI-compatible storage provisioner like Longhorn or Rook-Ceph
  • Deploy an ingress controller (nginx-ingress or Traefik) and test external routing
  • Set resource requests and limits for every production workload
  • Monitor CPU throttling metrics with kubectl top and cAdvisor
  • Tune etcd storage backend with --quota-backend-bytes and dedicated disk
  • Enable Prometheus node-exporter on every node for performance visibility

FAQ

How much RAM and CPU does a single-node Kubernetes cluster need?

A single control-plane node requires at least 2 CPU cores and 2GB RAM for kubeadm, kubelet, and a small CNI plugin like Flannel. Production clusters need 4 CPU and 8GB per control-plane node to handle etcd, API server, scheduler, and controller-manager under load. Worker nodes start at 2 CPU and 2GB but scale based on workload—reserve 10-15% overhead for system pods like kube-proxy and CNI agents.

Why do pods start slowly even when the node has free CPU?

CPU throttling from cgroups v1 default limits is the usual culprit. The kernel enforces cpu.cfs_quota_us even when cores are idle, causing kubelet and container runtime to wait for quota refills. Switch to cgroups v2, set systemd as the cgroup driver in containerd and kubelet configs, and raise cpu.max values for high-churn namespaces. Also check image pull time—large images over slow registries add 20-60 seconds per pod start.

Which CNI plugin performs best on bare metal?

Cilium with eBPF bypasses iptables and delivers sub-0.2ms pod-to-pod latency on modern kernels (5.10+). Calico is a solid middle ground with good policy support and 0.5-1ms overhead. Flannel VXLAN is simplest to deploy but adds 1-2ms encapsulation latency. For raw throughput on 10GbE or faster links, Cilium or Calico in direct-routing mode (no overlay) wins—measure with iperf3 between pods to confirm your hardware benefits.