Kubernetes Tutorial 2025: How to Deploy Your First Cluster: Practical Guide
Step-by-step Kubernetes tutorial for deploying your first cluster. Learn kubectl basics, pod management, and production-ready configurations with testable

On this page
- What Is Kubernetes and Why Use It
- Kubernetes Architecture: Control Plane and Worker Nodes
- Choosing Your Deployment Environment: Local vs Production
- Installing kubectl and Setting Up Minikube
- Deploying Your First Application with kubectl
- Writing Declarative YAML Manifests
- Scaling and Updating Applications
- Production Readiness: Health Checks and Resource Management
TL;DR — Key takeaways
- Kubernetes orchestrates containers across multiple nodes, automatically handling scaling, load balancing, and self-healing without manual intervention.
- A production-ready cluster requires at least three control plane nodes for high availability and separate worker nodes for application workloads.
- Start with Minikube or kind for local testing before deploying to cloud providers, allowing you to validate configurations without infrastructure costs.
- Use kubectl apply with YAML manifests instead of imperative commands to maintain version-controlled, reproducible deployments.
- Always set resource limits and requests in pod specifications to prevent resource exhaustion and ensure predictable application performance.
Kubernetes has become the standard for container orchestration, but deploying your first cluster can feel overwhelming. This tutorial walks you through the fundamentals: what Kubernetes does, how its components work together, and the exact steps to deploy a working cluster.
You'll learn the difference between local development clusters and production deployments, understand core Kubernetes resources like pods and services, and deploy your first application with testable commands. Whether you're a support engineer troubleshooting customer environments or a developer exploring infrastructure automation, this guide provides the foundation you need.
What Is Kubernetes and Why Use It
Kubernetes (often abbreviated as K8s) is an open-source container orchestration platform originally developed by Google. It automates the deployment, scaling, and management of containerized applications across clusters of machines.
Traditional deployments require manual server provisioning, load balancer configuration, and application monitoring. Kubernetes abstracts these tasks into declarative configurations. You define what you want (three replicas of an application, accessible via a load-balanced service), and Kubernetes continuously works to maintain that state.
The platform monitors application health, automatically restarts failed containers, distributes traffic across healthy instances, and scales resources based on demand. For hosting environments, this means reduced downtime, predictable resource usage, and faster deployment cycles.
- Self-healing: Automatically restarts containers that fail health checks
- Horizontal scaling: Adds or removes pod replicas based on CPU or custom metrics
- Service discovery: Provides DNS-based discovery and load balancing between pods
- Rolling updates: Deploys new versions without downtime using gradual rollout strategies
- Resource management: Enforces CPU and memory limits to prevent resource contention
Kubernetes Architecture: Control Plane and Worker Nodes
A Kubernetes cluster consists of two main components: the control plane and worker nodes. The control plane manages cluster state and scheduling decisions, while worker nodes run your application containers.
The control plane includes the API server (the entry point for all cluster operations), etcd (a distributed key-value store for cluster data), the scheduler (assigns pods to nodes), and controller managers (maintain desired state). In production, you run multiple control plane instances across different availability zones for fault tolerance.
Worker nodes run the kubelet agent, which communicates with the control plane and manages containers on that node. Each worker also runs kube-proxy for network routing and a container runtime (containerd or CRI-O) to execute containers. When you deploy an application, the scheduler assigns it to a worker node based on available resources and constraints.
- Control plane runs management components; worker nodes run application workloads
- etcd stores all cluster configuration and state; back it up before major changes
- kubelet on each node ensures containers match the desired pod specifications
- kube-proxy maintains network rules for service-to-pod communication
Choosing Your Deployment Environment: Local vs Production
Before deploying a cluster, decide between local development environments and production infrastructure. Local tools like Minikube, kind (Kubernetes in Docker), and k3d run single-node or multi-node clusters on your workstation. These are perfect for learning, testing manifests, and validating configurations before deploying to production.
For production deployments, you have three main paths: managed Kubernetes services (GKE, EKS, AKS), self-managed clusters on virtual machines, or bare-metal installations. Managed services handle control plane maintenance, upgrades, and backups, but cost more and offer less customization. Self-managed clusters give full control but require expertise in high availability, certificate management, and security patching.
This tutorial uses Minikube for initial learning because it requires minimal setup and runs entirely on your local machine. Once you understand the core concepts, you can apply the same kubectl commands and YAML manifests to production clusters.
- Minikube: Best for single-node local testing; includes dashboard and ingress addons
- kind: Multi-node clusters in Docker; faster startup than Minikube
- k3d: Lightweight k3s distribution; minimal resource overhead
- Managed services: Recommended for production unless you need infrastructure-level control
Installing kubectl and Setting Up Minikube
kubectl is the command-line tool for interacting with Kubernetes clusters. Install it before setting up your cluster. On Linux, download the binary directly:
curl -LO "https://dl.k8s.io/release/$(curl -L -s https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl" sudo install -o root -g root -m 0755 kubectl /usr/local/bin/kubectl
For macOS, use Homebrew: brew install kubectl. On Windows, download the executable from the official Kubernetes release page and add it to your PATH.
Next, install Minikube. On Linux:
curl -LO https://storage.googleapis.com/minikube/releases/latest/minikube-linux-amd64 sudo install minikube-linux-amd64 /usr/local/bin/minikube
macOS users can use brew install minikube. Windows users should download the installer from the Minikube releases page.
Start your first cluster with minikube start. This creates a single-node cluster running in a virtual machine or Docker container (depending on your system's driver). The process downloads the Kubernetes binaries, starts the control plane components, and configures kubectl to connect to the new cluster.
Verify the installation with kubectl cluster-info. You should see the control plane endpoint. Run kubectl get nodes to confirm the node is ready.
Deploying Your First Application with kubectl
Kubernetes runs applications in pods, the smallest deployable units. A pod contains one or more containers that share networking and storage. In production, you rarely create pods directly; instead, you use Deployments, which manage pod replicas and handle rolling updates.
Create a deployment using kubectl create deployment nginx --image=nginx:latest. This creates a Deployment resource that maintains one replica of an nginx container. Check the deployment status with kubectl get deployments and kubectl get pods.
The pod takes 10-30 seconds to start as Kubernetes pulls the container image. Use kubectl describe pod <pod-name> to see detailed events if the pod remains in Pending or ImagePullBackOff status.
Expose the deployment as a service to make it accessible: kubectl expose deployment nginx --type=NodePort --port=80. This creates a Service resource that load-balances traffic to the nginx pods. NodePort services expose applications on a port on each cluster node (30000-32767 range).
For Minikube, access the service with minikube service nginx. This opens the service URL in your browser. For production clusters, you'd use LoadBalancer or Ingress resources instead of NodePort.
- kubectl create: Quick imperative deployments; useful for testing
- kubectl apply: Declarative deployments using YAML files; recommended for production
- kubectl get: List resources; add -o wide for additional columns
- kubectl describe: Detailed resource information including events
- kubectl logs <pod-name>: View container output for troubleshooting
Writing Declarative YAML Manifests
Imperative commands like kubectl create are useful for experimentation, but production deployments require version-controlled, reproducible configurations. YAML manifests define your desired cluster state declaratively.
Here's a basic deployment manifest:
apiVersion: apps/v1 kind: Deployment metadata: name: nginx-deployment spec: replicas: 3 selector: matchLabels: app: nginx template: metadata: labels: app: nginx spec: containers: - name: nginx image: nginx:1.25 ports: - containerPort: 80 resources: requests: memory: "64Mi" cpu: "100m" limits: memory: "128Mi" cpu: "200m"
Save this as deployment.yaml and apply it with kubectl apply -f deployment.yaml. The spec.replicas field tells Kubernetes to maintain three identical pods. The selector.matchLabels connects the Deployment to its pods using label matching.
Always set resource requests and limits. Requests define the minimum resources guaranteed to the pod; limits define the maximum it can consume. Without these, a misbehaving pod can starve other applications of CPU or memory.
Create a matching service manifest:
apiVersion: v1 kind: Service metadata: name: nginx-service spec: selector: app: nginx ports: - protocol: TCP port: 80 targetPort: 80 type: ClusterIP
Apply with kubectl apply -f service.yaml. ClusterIP services are only accessible within the cluster. For external access, change type to LoadBalancer (on cloud providers) or use an Ingress controller.
Scaling and Updating Applications
Kubernetes makes scaling straightforward. To scale the nginx deployment to five replicas, use kubectl scale deployment nginx-deployment --replicas=5. Kubernetes immediately starts two additional pods on available worker nodes.
For production, configure the Horizontal Pod Autoscaler (HPA) to scale automatically based on metrics: kubectl autoscale deployment nginx-deployment --cpu-percent=70 --min=3 --max=10. This maintains CPU utilization around 70% by adding or removing replicas.
Update the application by changing the image version in your manifest: change image: nginx:1.25 to image: nginx:1.26, then run kubectl apply -f deployment.yaml. Kubernetes performs a rolling update, gradually replacing old pods with new ones while maintaining availability.
Monitor the rollout with kubectl rollout status deployment/nginx-deployment. If issues occur, rollback immediately: kubectl rollout undo deployment/nginx-deployment. This reverts to the previous working version.
- Scale manually with kubectl scale or declaratively by editing replicas in YAML
- Use HPA for automatic scaling based on CPU, memory, or custom metrics
- Rolling updates replace pods gradually; configure maxSurge and maxUnavailable for control
- Always test updates in a staging environment before applying to production
Production Readiness: Health Checks and Resource Management
Production deployments require health checks to ensure Kubernetes routes traffic only to healthy pods. Add liveness and readiness probes to your container spec:
livenessProbe: httpGet: path: /healthz port: 80 initialDelaySeconds: 30 periodSeconds: 10 readinessProbe: httpGet: path: /ready port: 80 initialDelaySeconds: 5 periodSeconds: 5
Liveness probes detect containers that are running but unable to make progress (deadlocked or hung). Kubernetes restarts pods that fail liveness checks. Readiness probes determine when a pod is ready to accept traffic; pods failing readiness checks are removed from service endpoints until they recover.
Set appropriate resource requests and limits based on application profiling. Under-provisioning causes pod eviction under memory pressure; over-provisioning wastes cluster capacity. Monitor actual resource usage with kubectl top pods after deploying.
Use namespaces to isolate environments and teams: kubectl create namespace production. Deploy resources to specific namespaces with kubectl apply -f deployment.yaml -n production. This prevents accidental changes across environments and enables namespace-level resource quotas.
Quick troubleshooting checklist
- Install kubectl and verify version with kubectl version --client
- Install Minikube or choose a production cluster environment
- Start your cluster with minikube start or connect to existing cluster
- Verify cluster connectivity: kubectl cluster-info and kubectl get nodes
- Create a test deployment: kubectl create deployment nginx --image=nginx
- Check pod status: kubectl get pods -w (watch mode)
- Expose the deployment: kubectl expose deployment nginx --type=NodePort --port=80
- Access the service: minikube service nginx (Minikube) or kubectl get svc (production)
- Write YAML manifests with resource limits for production deployments
- Add liveness and readiness probes to detect unhealthy pods
- Test scaling: kubectl scale deployment nginx --replicas=3
- Configure HPA for automatic scaling based on metrics
- Test rolling updates: change image version and kubectl apply
- Practice rollback: kubectl rollout undo deployment/nginx
- Use namespaces to separate environments: kubectl create namespace staging
FAQ
What is the difference between a pod and a deployment in Kubernetes?
A pod is the smallest deployable unit in Kubernetes containing one or more containers that share storage and networking. A Deployment is a higher-level resource that manages pod replicas, handles rolling updates, and ensures the desired number of pods are always running. Deployments automatically recreate pods that fail and provide declarative update strategies, while pods are ephemeral and not automatically replaced when deleted.
How much resources do I need to run a Kubernetes cluster?
Minikube requires at least 2 GB RAM and 2 CPU cores for basic testing. Production clusters need a minimum of 3 control plane nodes (2 GB RAM, 2 CPUs each) for high availability and separate worker nodes sized according to your workload. A small production cluster typically starts with 3 control plane nodes and 3-5 worker nodes with 4-8 GB RAM and 2-4 CPUs each, plus storage for persistent volumes.
Why is my pod stuck in Pending status?
Pods remain in Pending status when Kubernetes cannot schedule them to a node. Common causes include insufficient cluster resources (CPU or memory), unmet node selector requirements, or missing persistent volume claims. Run kubectl describe pod <pod-name> to see scheduling events and error messages. Check available node resources with kubectl top nodes and verify your pod's resource requests are not higher than available capacity.
Should I use kubectl create or kubectl apply for deployments?
Use kubectl apply with YAML manifests for production deployments because it supports version control, reproducible configurations, and incremental updates. kubectl create is useful for quick testing and imperative commands but does not track configuration history. kubectl apply compares your manifest against the cluster state and only applies changes, making it safer for production use and enabling GitOps workflows.
How do I access applications running in my Kubernetes cluster?
Access methods depend on your service type and environment. ClusterIP services are only accessible from within the cluster. NodePort services expose applications on a port (30000-32767) on every node. LoadBalancer services automatically provision cloud load balancers for external access. For Minikube, use minikube service <service-name> to access NodePort services. In production, use Ingress controllers for HTTP/HTTPS routing or LoadBalancer services for TCP/UDP traffic.
Related articles
- Hosting OperationsSelf-Hosted App Deployment Fails? Check DNS, SSL, Reverse Proxy, and Logs FirstTroubleshoot failed self-hosted app deployments by checking DNS, SSL, reverse proxy routing, container status, logs, and ports.
- Hosting OperationsSelf-Hosted PaaS on a VPS: What to Check Before Installing Coolify, Dokploy, or CapRoverA hosting support checklist for preparing a VPS before installing self-hosted PaaS tools like Coolify, Dokploy, or CapRover.
- Hosting OperationsLinux Server Security Lessons from the Arch Linux Malware Package IncidentPractical Linux server security checklist for VPS admins after package malware concerns, with safe checks, rollback steps, and support guidance.
- Hosting OperationsAWS Lightsail Hong Kong VPS Latency: Practical Hosting Guide for IndonesiaLearn how to test AWS Lightsail Hong Kong VPS latency, compare regions, migrate safely, and troubleshoot hosting performance.
- Hosting OperationsCloudflare Tomorrow Watchlist: A Practical Hosting Operations GuidePractical Cloudflare troubleshooting checklist for DNS, SSL, caching, WAF, origin health, safe testing, and rollback planning.
- Hosting OperationsNetwork Safety Checklist for AI Agent Skills in Hosting OperationsAudit AI agent skills safely with network checks, secret protection, sandbox testing, rollback steps, and hosting support troubleshooting guidance.