← Back to Engineering Insights

Cloud Infrastructure & Kubernetes

Production Container Orchestration: Docker & Kubernetes Best Practices

Key Architecture Takeaways

  • Always Define Resource Requests and Limits: Unbounded pods cause noisy-neighbor node CPU starvation and unexpected Out-Of-Memory (OOMKilled) cluster evictions.
  • Differentiate Readiness from Liveness Probes: Liveness probes restart failing containers; Readiness probes control ingress routing. Flapping services should fail readiness without triggering catastrophic reboot loops.
  • Implement Graceful Pod Shutdown: Handle SIGTERM signals cleanly with application draining timeouts and preStop lifecycle sleep hooks to prevent dropped inflight HTTP connections.
  • Enforce Security Contexts: Block privilege escalation, enforce read-only root filesystems, drop all Linux capabilities, and run exclusively as unprivileged user IDs.

Running containers in production is fundamentally different from running docker run on a local development laptop. In production enterprise environments, systems must automatically self-heal after hardware node crashes, dynamically scale during traffic surges, and roll out zero-downtime updates without dropping a single active customer transaction.

Kubernetes has emerged as the definitive operating system of the modern cloud. Yet, misconfigured Kubernetes manifests and poorly optimized container images frequently introduce instability, cascading cluster outages, and inflated cloud bills. In this comprehensive guide, we examine proven operational blueprints for architecting production-grade containerized workloads on Kubernetes.

Advertisement

1. The Anatomy of a Bulletproof Kubernetes Deployment Manifest

A production Kubernetes deployment must account for scheduling constraints, resource quotas, security controls, and rolling update strategies. The following manifest exemplifies an enterprise-hardened deployment configuration:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: enterprise-order-service
  namespace: production
  labels:
    app.kubernetes.io/name: order-service
    app.kubernetes.io/part-of: ecommerce-platform
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1        # Spin up 1 new pod before tearing down old ones
      maxUnavailable: 0  # Guarantee zero downtime during release
  selector:
    matchLabels:
      app: order-service
  template:
    metadata:
      labels:
        app: order-service
    spec:
      # Prevent all pods from landing on the same physical worker node
      affinity:
        podAntiAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
            - weight: 100
              podAffinityTerm:
                labelSelector:
                  matchExpressions:
                    - key: app
                      operator: In
                      values: [order-service]
                topologyKey: kubernetes.io/hostname
      securityContext:
        runAsNonRoot: true
        runAsUser: 10001
        runAsGroup: 10001
        fsGroup: 10001
      containers:
        - name: order-service
          image: ghcr.io/sunsmitsoftware/order-service:v2.4.1
          imagePullPolicy: IfNotPresent
          ports:
            - containerPort: 8080
              name: http
          resources:
            requests:
              cpu: "250m"
              memory: "512Mi"
            limits:
              cpu: "1000m"
              memory: "1Gi"
          lifecycle:
            preStop:
              exec:
                command: ["/bin/sh", "-c", "sleep 15"] # Grace period for ingress drain
          readinessProbe:
            httpGet:
              path: /health/ready
              port: 8080
            initialDelaySeconds: 5
            periodSeconds: 10
            timeoutSeconds: 2
            failureThreshold: 3
          livenessProbe:
            httpGet:
              path: /health/live
              port: 8080
            initialDelaySeconds: 15
            periodSeconds: 20
            timeoutSeconds: 3
            failureThreshold: 3

2. Master Probe Configuration: Startup vs Liveness vs Readiness

Health checks are the core mechanism through which Kubernetes maintains cluster availability. However, conflating probe types leads to severe production incidents:

Probe Type Primary Purpose Failure Action by Kubelet What to Check Inside Application
Startup Probe Guards slow-starting applications (JVM warm-up, DB migrations) Kills and restarts pod only after entire startup timeout expires Checks if initial configuration and dependency boot completed
Readiness Probe Determines if pod should receive traffic from Ingress/Service Removes pod IP from Service endpoints; does not restart pod Database connectivity, cache availability, internal thread pools
Liveness Probe Detects deadlocks or unrecoverable frozen processes Immediately terminates container and schedules fresh reboot Checks internal process responsiveness (e.g., pinging an internal loop)

Common Anti-Pattern: Placing external database connection checks inside a Liveness Probe. If your PostgreSQL database suffers transient failover for 30 seconds, all 50 service pods will fail their liveness checks simultaneously, causing Kubelet to restart every pod at the same time and triggering a catastrophic Thundering Herd Outage. Database checks belong exclusively in the Readiness Probe!

3. Zero-Downtime Graceful Termination: Handling SIGTERM

When Kubernetes scales down a deployment or rolls out an update, it sends a SIGTERM signal to the container process. However, cluster networking (iptables, Ingress, kube-proxy) requires several seconds to propagate pod IP removal across all cluster nodes. If an application exits immediately upon receiving SIGTERM, inflight HTTP requests arriving from the edge gateway will encounter 502 Bad Gateway or connection reset errors.

To achieve truly zero-downtime deployments:

  1. Add a preStop lifecycle hook that executes a sleep 10-15 command before sending the signal to the binary. This gives cluster ingress routers time to stop sending new traffic to the terminating pod.
  2. Configure your web framework (e.g., ASP.NET Core, Express, Go net/http) to allow active inflight requests 15–30 seconds to complete processing before forcibly shutting down.
  3. Ensure terminationGracePeriodSeconds on the pod manifest exceeds the combined preStop sleep and application drain timeout (typically set to 45–60 seconds).

4. Horizontal Pod Autoscaling (HPA) with Custom Metrics

Relying solely on CPU utilization to trigger autoscaling is inadequate for I/O-bound or message-driven services. A service consuming 15% CPU might be backlogged by 500,000 pending orders in a message queue.

Enterprise clusters deploy KEDA (Kubernetes Event-driven Autoscaling) or the Prometheus Adapter to autoscale workloads based on business metrics, such as Kafka consumer lag, RabbitMQ queue depth, or incoming HTTP request rate (RPS), scaling pods proactively before response latencies degrade.

AP
Ashu Patel

Lead Solutions Architect at Sunsmit Software. Ashu specializes in Kubernetes cloud architecture, production container hardening, DevOps automation, and microservices reliability engineering.