Kubernetes Security Best Practices: 9 Controls for 2026
Harden your K8s clusters with RBAC, network policies, pod security standards, and secrets management. Nine production-ready controls.

On this page
- Role-Based Access Control: Enforce Least Privilege
- Network Policies: Default-Deny and Explicit Allow
- Pod Security Standards: Enforce Baseline and Restricted Modes
- Secrets Management: Move to External Stores
- Image Security and Supply-Chain Controls
- Audit Logging and Monitoring
- Resource Limits and Admission Policies
- Secure etcd and API Server Access
- Testing and Validation in Staging
TL;DR — Key takeaways
- Role-based access control (RBAC) should enforce least-privilege policies with namespace-scoped roles instead of cluster-wide bindings.
- Network policies isolate pods by default and allow only explicitly permitted traffic, reducing lateral movement after a breach.
- Pod Security Standards replace deprecated PodSecurityPolicies and enforce baseline security requirements at the namespace level.
- External secrets managers (Vault, cloud KMS) prevent plaintext secrets in Git and provide audit logs for every secret access.
- Supply-chain controls like image signing and admission webhooks block unverified container images before they run.
Kubernetes clusters are often the weakest link in infrastructure security. Default configurations leave RBAC too open, pods running as root, and secrets stored in plaintext. A single misconfigured role binding can expose your entire cluster.
This guide walks through nine security controls that harden production clusters without breaking existing workloads. Each section identifies the bottleneck—overly permissive access, unfiltered network traffic, or insecure defaults—and provides tuning steps with specific kubectl commands and YAML examples. These are the configurations I verify when troubleshooting compromised environments.
Role-Based Access Control: Enforce Least Privilege
RBAC controls who can read, modify, or delete resources in your cluster. The default ServiceAccount in each namespace has no permissions, but many tutorials tell you to create a ClusterRole with broad verbs and bind it cluster-wide. That's the bottleneck.
Start by auditing existing bindings. Run kubectl get clusterrolebindings -o json and look for bindings that grant cluster-admin or use wildcards in resources or verbs. In support tickets I handled, the usual culprit was a deployment script that created a ClusterRoleBinding to simplify setup, then never removed it.
Replace cluster-wide bindings with namespace-scoped Roles. If a service account only needs to read ConfigMaps in the 'app' namespace, create a Role in that namespace with get and list verbs on configmaps, then bind it with a RoleBinding. Never use ClusterRoles unless a component truly needs cluster-wide visibility, like a monitoring agent or cluster autoscaler.
Test each role change in a non-production namespace first. Deploy your workload, exec into a pod, and run kubectl auth can-i get pods to confirm the service account has exactly the permissions it needs. If the pod starts but can't read required resources, add specific verbs one at a time until it works. Document the minimum required permissions in your deployment README.
Network Policies: Default-Deny and Explicit Allow
Without network policies, every pod can reach every other pod on any port. An attacker who compromises one container can scan your entire cluster and pivot to databases or internal APIs.
Implement a default-deny policy in each namespace. This two-line policy blocks all ingress and egress traffic unless another policy explicitly allows it:
After applying default-deny, your existing pods will lose connectivity. That's expected. Now create targeted allow rules for each service. A web frontend needs egress to the backend API on port 8080 and ingress from the ingress controller. The backend needs egress to the database on port 5432. Write one NetworkPolicy per service with podSelector and ports specified.
Test connectivity with a debug pod. Run kubectl run netshoot --rm -i --tty --image nicolaka/netshoot -- /bin/bash, then use curl or nc to verify allowed paths work and denied paths fail. Check your CNI plugin's logs if policies aren't enforced; some clusters need the network policy feature gate explicitly enabled.
- apiVersion: networking.k8s.io/v1
- kind: NetworkPolicy
- metadata:
- name: default-deny-all
- namespace: production
- spec:
- podSelector: {}
- policyTypes:
- - Ingress
- - Egress
Pod Security Standards: Enforce Baseline and Restricted Modes
Pod Security Standards define three levels: privileged (unrestricted), baseline (blocks known privilege escalations), and restricted (defense-in-depth for critical workloads). Most clusters run everything in privileged mode because that's the default.
Enable Pod Security admission at the namespace level. Add labels to your namespace to enforce a policy mode. For production workloads, start with baseline in audit mode to see what would be blocked without actually rejecting pods. Check the audit logs for violations, fix your pod specs, then switch to enforce mode.
Restricted mode is the target for production. It requires non-root users, drops all capabilities, sets readOnlyRootFilesystem, and disallows privilege escalation. Many container images fail restricted mode because they write to /tmp or need specific capabilities. You'll need to rebuild images with a non-root USER directive and explicit volume mounts for writable paths.
Before enforcing restricted mode cluster-wide, test each deployment individually. Create a separate namespace with restricted mode enforced, deploy your workload there, and run integration tests. If a pod fails to start, check the events with kubectl describe pod. The error message will tell you which security field caused the rejection.
- kubectl label namespace production pod-security.kubernetes.io/enforce=baseline
- kubectl label namespace production pod-security.kubernetes.io/audit=restricted
- kubectl label namespace production pod-security.kubernetes.io/warn=restricted
Secrets Management: Move to External Stores
Kubernetes Secrets are base64-encoded, not encrypted, and stored in etcd. Anyone with etcd access or permission to read secrets in a namespace can decode them instantly. Git repositories full of manifests with hardcoded secrets are a daily occurrence in breach reports.
Use an external secrets operator to sync secrets from a real secrets manager. HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, and Google Secret Manager all have Kubernetes operators that create Secret objects from external sources on demand. The secret value never touches your Git repository, and every access is logged.
If you can't deploy an external operator immediately, enable encryption at rest for etcd. Add an EncryptionConfiguration to your API server with a provider like aescbc or kms. Existing secrets won't be encrypted until you rewrite them—run kubectl get secrets --all-namespaces -o json | kubectl replace -f - to force a rewrite. Rotate the encryption key every 90 days.
Service account tokens are secrets too. Kubernetes 1.24 stopped auto-generating long-lived tokens and switched to time-bound tokens projected into pods. If you're on an older version, delete unused service accounts and rotate tokens for the ones you keep. Check for tokens that haven't been used in six months and remove them.
Image Security and Supply-Chain Controls
Container images from public registries can contain malware, backdoors, or vulnerable libraries. An admission webhook that validates image signatures and scans for CVEs blocks malicious images before they run.
Sign your images with cosign or Notary. After building an image, generate a signature with your private key and push the signature to the registry alongside the image. Deploy an admission controller like Kyverno or OPA Gatekeeper with a policy that rejects pods using unsigned images. This prevents an attacker who compromises your registry from injecting a backdoored image.
Scan images during CI and at runtime. Trivy, Grype, and Clair can detect known CVEs in your base images and dependencies. Integrate a scanner into your build pipeline to fail builds with high-severity vulnerabilities. Run the same scanner as a periodic job in your cluster to catch newly disclosed CVEs in running images.
Set ImagePullPolicy to Always and use immutable tags. Never use 'latest' or rolling tags like 'v1'—they hide what version is actually deployed and make rollbacks impossible. Use SHA digests in production manifests (image: myapp@sha256:abc123...) to guarantee you're running the exact build you tested.
Audit Logging and Monitoring
Audit logs record every API request to your cluster—who did what, when, and whether it succeeded. Without audit logs, you can't investigate a breach or prove compliance.
Enable audit logging on your API server with an audit policy file. Start with a basic policy that logs metadata for all requests and request/response bodies for sensitive resources like secrets, configmaps, and roles. Be careful with the response body level—it will log secret values if someone does kubectl get secret -o yaml.
Ship logs to a separate system outside the cluster. If your cluster is compromised, an attacker with admin access can delete audit logs stored in the same environment. Use Fluentd or Promtail to forward logs to Loki, Elasticsearch, or a cloud logging service. Retain logs for at least 90 days to meet most compliance requirements.
Monitor for anomalies. Set alerts for kubectl exec commands, changes to RBAC policies, secret reads by unexpected service accounts, and repeated authentication failures. In production incidents, the first sign of trouble was often a spike in failed API calls from a compromised token.
Resource Limits and Admission Policies
Pods without resource limits can consume all CPU and memory on a node, taking down other workloads. This is a denial-of-service risk, not just a performance issue.
Set requests and limits on every container. Requests reserve resources; limits cap usage. A container without limits can burst to use an entire node's CPU during a traffic spike or a memory leak. Use LimitRanges to enforce defaults in namespaces where developers forget to set them.
Use ResourceQuotas to cap total usage per namespace. If a team deploys 50 replicas by mistake, a quota will reject the excess pods instead of taking down the cluster. Set quotas for CPU, memory, persistent volume claims, and pod counts based on your team's normal usage plus a safety margin.
Test resource constraints before rolling them out. Deploy a load generator and gradually increase traffic until you hit the limit. The pod should throttle gracefully, not crash. If a container restarts with an OOMKilled status, your memory limit is too low. Increase it in 50MB increments until the pod survives peak load.
Secure etcd and API Server Access
etcd stores your cluster state, including secrets and configuration. If someone gains direct access to etcd, RBAC and network policies don't matter—they can read or modify anything.
Run etcd on dedicated nodes that aren't accessible from worker nodes. Use mutual TLS for all etcd communication and rotate certificates yearly. Set --client-cert-auth=true and --peer-client-cert-auth=true to require valid certificates. Never expose etcd's client port to the internet or to namespaces where users run pods.
Restrict API server access with an authentication webhook or OIDC integration. Username/password auth is deprecated; switch to token-based or certificate-based authentication tied to your identity provider. Use Node authorizer mode so kubelets can only modify resources on their own node.
Enable anonymous-auth=false unless you have a specific use case for unauthenticated requests. In older clusters, anonymous users had discovery permissions by default. Check with kubectl get clusterrolebinding system:discovery and remove it if you don't need unauthenticated access to API discovery endpoints.
Testing and Validation in Staging
Security controls that break production deployments get disabled immediately. Test every change in a staging cluster with realistic workloads before enforcing it.
Create a staging namespace that mirrors production configuration—same resource limits, same network policies, same pod security mode. Deploy your application stack and run integration tests. If tests pass in staging with restricted security controls, you can roll them out to production with confidence.
Use kubectl dry-run for manifest changes. Before applying a new NetworkPolicy or RBAC rule, run kubectl apply --dry-run=server -f policy.yaml to see if the API server would accept it. This catches typos and invalid selectors before they disrupt traffic.
Document rollback procedures for every control. Network policies can be removed instantly with kubectl delete networkpolicy, but reverting pod security modes requires removing namespace labels and restarting pods. Write a runbook with exact commands so anyone on-call can revert a breaking change in under five minutes.
Quick troubleshooting checklist
- Audit existing RBAC bindings and remove cluster-admin where unnecessary
- Enable default-deny network policies in every namespace
- Set pod security admission to 'restricted' mode for production namespaces
- Migrate Secrets to an external secrets operator or cloud KMS
- Configure image pull policies to reject unsigned images
- Enable audit logging with a 90-day retention policy
- Set resource limits on every pod to prevent resource exhaustion attacks
- Review and rotate service account tokens quarterly
- Test security controls in a staging cluster before production rollout
FAQ
What is the difference between pod security policies and pod security standards?
PodSecurityPolicies were deprecated in Kubernetes 1.21 and removed in 1.25. Pod Security Standards are the replacement, enforced by the Pod Security admission controller at the namespace level with three modes: privileged, baseline, and restricted. They require less configuration and integrate directly into the API server.
How do I know if my RBAC configuration is too permissive?
Run kubectl auth can-i --list --as=system:serviceaccount:namespace:serviceaccount-name to see what permissions a service account has. If you see cluster-wide verbs like create, delete, or * on core resources, investigate whether namespace-scoped roles can replace them. Tools like rbac-lookup and krane can audit your cluster for overly broad bindings.
Do network policies affect cluster performance?
Network policies add minimal overhead when using modern CNI plugins like Calico or Cilium, typically under 5% latency increase. The performance impact comes from connection tracking, not packet filtering. Test with realistic traffic patterns in staging before rolling out default-deny policies to production.
Related articles
- Hosting OperationsSelf-Hosted App Deployment Fails? Check DNS, SSL, Reverse Proxy, and Logs FirstTroubleshoot failed self-hosted app deployments by checking DNS, SSL, reverse proxy routing, container status, logs, and ports.
- Hosting OperationsSelf-Hosted PaaS on a VPS: What to Check Before Installing Coolify, Dokploy, or CapRoverA hosting support checklist for preparing a VPS before installing self-hosted PaaS tools like Coolify, Dokploy, or CapRover.
- Hosting OperationsLinux Server Security Lessons from the Arch Linux Malware Package IncidentPractical Linux server security checklist for VPS admins after package malware concerns, with safe checks, rollback steps, and support guidance.
- Hosting OperationsAWS Lightsail Hong Kong VPS Latency: Practical Hosting Guide for IndonesiaLearn how to test AWS Lightsail Hong Kong VPS latency, compare regions, migrate safely, and troubleshoot hosting performance.
- Hosting OperationsCloudflare Tomorrow Watchlist: A Practical Hosting Operations GuidePractical Cloudflare troubleshooting checklist for DNS, SSL, caching, WAF, origin health, safe testing, and rollback planning.
- Hosting OperationsNetwork Safety Checklist for AI Agent Skills in Hosting OperationsAudit AI agent skills safely with network checks, secret protection, sandbox testing, rollback steps, and hosting support troubleshooting guidance.