Cloud Infrastructure & Kubernetes
Production Container Orchestration: Docker & Kubernetes Best Practices
Key Architecture Takeaways
- Always Define Resource Requests and Limits: Unbounded pods cause noisy-neighbor node CPU starvation and unexpected Out-Of-Memory (OOMKilled) cluster evictions.
- Differentiate Readiness from Liveness Probes: Liveness probes restart failing containers; Readiness probes control ingress routing. Flapping services should fail readiness without triggering catastrophic reboot loops.
- Implement Graceful Pod Shutdown: Handle
SIGTERMsignals cleanly with application draining timeouts andpreStoplifecycle sleep hooks to prevent dropped inflight HTTP connections. - Enforce Security Contexts: Block privilege escalation, enforce read-only root filesystems, drop all Linux capabilities, and run exclusively as unprivileged user IDs.
Running containers in production is fundamentally different from running docker run on a local development laptop. In production enterprise environments, systems must automatically self-heal after hardware node crashes, dynamically scale during traffic surges, and roll out zero-downtime updates without dropping a single active customer transaction.
Kubernetes has emerged as the definitive operating system of the modern cloud. Yet, misconfigured Kubernetes manifests and poorly optimized container images frequently introduce instability, cascading cluster outages, and inflated cloud bills. In this comprehensive guide, we examine proven operational blueprints for architecting production-grade containerized workloads on Kubernetes.
1. The Anatomy of a Bulletproof Kubernetes Deployment Manifest
A production Kubernetes deployment must account for scheduling constraints, resource quotas, security controls, and rolling update strategies. The following manifest exemplifies an enterprise-hardened deployment configuration:
apiVersion: apps/v1
kind: Deployment
metadata:
name: enterprise-order-service
namespace: production
labels:
app.kubernetes.io/name: order-service
app.kubernetes.io/part-of: ecommerce-platform
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1 # Spin up 1 new pod before tearing down old ones
maxUnavailable: 0 # Guarantee zero downtime during release
selector:
matchLabels:
app: order-service
template:
metadata:
labels:
app: order-service
spec:
# Prevent all pods from landing on the same physical worker node
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchExpressions:
- key: app
operator: In
values: [order-service]
topologyKey: kubernetes.io/hostname
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
containers:
- name: order-service
image: ghcr.io/sunsmitsoftware/order-service:v2.4.1
imagePullPolicy: IfNotPresent
ports:
- containerPort: 8080
name: http
resources:
requests:
cpu: "250m"
memory: "512Mi"
limits:
cpu: "1000m"
memory: "1Gi"
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 15"] # Grace period for ingress drain
readinessProbe:
httpGet:
path: /health/ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
livenessProbe:
httpGet:
path: /health/live
port: 8080
initialDelaySeconds: 15
periodSeconds: 20
timeoutSeconds: 3
failureThreshold: 3
2. Master Probe Configuration: Startup vs Liveness vs Readiness
Health checks are the core mechanism through which Kubernetes maintains cluster availability. However, conflating probe types leads to severe production incidents:
| Probe Type | Primary Purpose | Failure Action by Kubelet | What to Check Inside Application |
|---|---|---|---|
| Startup Probe | Guards slow-starting applications (JVM warm-up, DB migrations) | Kills and restarts pod only after entire startup timeout expires | Checks if initial configuration and dependency boot completed |
| Readiness Probe | Determines if pod should receive traffic from Ingress/Service | Removes pod IP from Service endpoints; does not restart pod | Database connectivity, cache availability, internal thread pools |
| Liveness Probe | Detects deadlocks or unrecoverable frozen processes | Immediately terminates container and schedules fresh reboot | Checks internal process responsiveness (e.g., pinging an internal loop) |
Common Anti-Pattern: Placing external database connection checks inside a Liveness Probe. If your PostgreSQL database suffers transient failover for 30 seconds, all 50 service pods will fail their liveness checks simultaneously, causing Kubelet to restart every pod at the same time and triggering a catastrophic Thundering Herd Outage. Database checks belong exclusively in the Readiness Probe!
3. Zero-Downtime Graceful Termination: Handling SIGTERM
When Kubernetes scales down a deployment or rolls out an update, it sends a SIGTERM signal to the container process. However, cluster networking (iptables, Ingress, kube-proxy) requires several seconds to propagate pod IP removal across all cluster nodes. If an application exits immediately upon receiving SIGTERM, inflight HTTP requests arriving from the edge gateway will encounter 502 Bad Gateway or connection reset errors.
To achieve truly zero-downtime deployments:
- Add a
preStoplifecycle hook that executes asleep 10-15command before sending the signal to the binary. This gives cluster ingress routers time to stop sending new traffic to the terminating pod. - Configure your web framework (e.g., ASP.NET Core, Express, Go net/http) to allow active inflight requests 15–30 seconds to complete processing before forcibly shutting down.
- Ensure
terminationGracePeriodSecondson the pod manifest exceeds the combined preStop sleep and application drain timeout (typically set to 45–60 seconds).
4. Horizontal Pod Autoscaling (HPA) with Custom Metrics
Relying solely on CPU utilization to trigger autoscaling is inadequate for I/O-bound or message-driven services. A service consuming 15% CPU might be backlogged by 500,000 pending orders in a message queue.
Enterprise clusters deploy KEDA (Kubernetes Event-driven Autoscaling) or the Prometheus Adapter to autoscale workloads based on business metrics, such as Kafka consumer lag, RabbitMQ queue depth, or incoming HTTP request rate (RPS), scaling pods proactively before response latencies degrade.