Kubernetes Core K8s Workloads

Kubernetes Cluster Architecture & Core Workloads: Pods, Deployments, and Scaling

⏱ 12 min read • Level: Intermediate • Updated: Sep 30, 2026

1. Executive Overview & Industry Context

Kubernetes has cemented its position as the universal operating system for cloud-native enterprise computing. Originally designed by Google based on fifteen years of production experience running internal Borg cluster managers, Kubernetes provides a declarative, container-centric platform for automating the deployment, scaling, networking, and lifecycle management of application workloads across distributed server fleets. By abstracting physical or virtual machine infrastructure into a unified compute pool, Kubernetes enables engineering organizations to deploy software with predictable velocity, resilience, and multi-cloud portability.

However, running production workloads on Kubernetes requires software engineers to transition from imperative infrastructure management to declarative reconciliation. Rather than manually provisioning servers or executing shell scripts to start processes, developers author desired-state manifests. The Kubernetes control plane continuously reconciles the observed state of the cluster with this declared desired state. Mastering cluster architecture, Pod lifecycle mechanics, and zero-downtime Deployment strategies is the foundational prerequisite for cloud-native DevOps engineering.

2. Core Learning Objectives

By concluding this technical module, cloud architects and Kubernetes practitioners will demonstrate verifiable competency in the following capabilities:

  • Control Plane & Worker Node Topology: Analyze the interactions between kube-apiserver, etcd, kube-scheduler, kube-controller-manager, and kubelet.
  • Pod Lifecycle & Spec Formulation: Author multi-container Pod manifests utilizing init containers, sidecars, and resource requests/limits.
  • Deployment Strategies: Configure RollingUpdate parameters (maxSurge, maxUnavailable) and execute zero-downtime rollbacks using kubectl rollout.
  • Horizontal Pod Autoscaling (HPA): Implement metrics-driven autoscaling based on CPU, memory, and custom Prometheus metrics.

3. Theoretical Foundations & Architecture

A Kubernetes cluster consists of two distinct operational planes: the Control Plane and Worker Nodes. The Control Plane components coordinate the entire cluster:

  • kube-apiserver: The central administrative gateway exposing the Kubernetes API; validates and configures data for pods, services, and controllers. All cluster communications route through the API server.
  • etcd: Highly available, distributed key-value store containing the complete source of truth and state of the cluster.
  • kube-scheduler: Monitors newly created Pods with no assigned node, selecting an optimal worker node based on resource requirements, affinity/anti-affinity rules, and taints/tolerations.
  • kube-controller-manager: Executes core reconciliation loops (Node Controller, Deployment Controller, EndpointSlice Controller, Job Controller) to drive actual state toward desired state.

On each Worker Node, two primary agents maintain container execution: the kubelet (node agent ensuring containers described in PodSpecs are running and healthy) and the kube-proxy (maintains network rules on nodes to allow network communication to Pods).

The Pod is the smallest deployable compute unit in Kubernetes, representing a shared execution context consisting of one or more tightly coupled containers sharing Linux network namespaces (IP address and port space) and storage volumes. Above Pods, higher-order controllers manage lifecycle concerns: ReplicaSets ensure a specified number of identical pod replicas are running, while Deployments provide declarative updates for Pods and ReplicaSets, orchestrating zero-downtime rolling updates.

4. Step-by-Step Implementation Guide & Code Demonstrations

The following production manifest demonstrates an enterprise Deployment with resource limits, liveness/readiness probes, RollingUpdate tuning, and an accompanying HorizontalPodAutoscaler (HPA):

# ==============================================================================
# PRODUCTION DEPLOYMENT MANIFEST: Transaction Ingress Microservice
# ==============================================================================
apiVersion: apps/v1
kind: Deployment
metadata:
  name: transaction-ingress
  namespace: production
  labels:
    app.kubernetes.io/name: transaction-ingress
    app.kubernetes.io/part-of: financial-ledger
spec:
  replicas: 3
  revisionHistoryLimit: 10
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 25%        # Allow temporary replica burst during deployment
      maxUnavailable: 0     # Guarantee 100% capacity availability during rollouts
  selector:
    matchLabels:
      app: transaction-ingress
  template:
    metadata:
      labels:
        app: transaction-ingress
    spec:
      containers:
        - name: ingress-service
          image: registry.enterprise.org/ledger/ingress:v2.4.0
          imagePullPolicy: IfNotPresent
          ports:
            - containerPort: 8080
              name: http
          resources:
            requests:
              cpu: "250m"       # Guaranteed reservation: 0.25 CPU cores
              memory: "512Mi"    # Guaranteed reservation: 512 Megabytes
            limits:
              cpu: "1000m"      # Hard ceiling: 1.0 CPU core (Throttled if exceeded)
              memory: "1Gi"      # Hard ceiling: 1.0 GiB (OOMKilled if exceeded)
          livenessProbe:
            httpGet:
              path: /healthz
              port: 8080
            initialDelaySeconds: 15
            periodSeconds: 10
            failureThreshold: 3
          readinessProbe:
            httpGet:
              path: /ready
              port: 8080
            initialDelaySeconds: 5
            periodSeconds: 5
            failureThreshold: 2

---
# ==============================================================================
# HORIZONTAL POD AUTOSCALER (HPA) v2
# ==============================================================================
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: transaction-ingress-hpa
  namespace: production
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: transaction-ingress
  minReplicas: 3
  maxReplicas: 15
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 75

Accompanying rollout inspection and rollback CLI commands:

# Monitor the rolling update in real-time
kubectl rollout status deployment/transaction-ingress -n production

# View historical deployment revisions
kubectl rollout history deployment/transaction-ingress -n production

# Execute instantaneous rollback to the previous revision if an anomaly occurs
kubectl rollout undo deployment/transaction-ingress -n production

5. Real-World Case Studies & Enterprise Production Scenarios

An e-commerce retail platform experienced intermittent 502 Bad Gateway outages during scheduled software releases. An architectural post-mortem revealed two critical configuration errors in their Deployment manifests: first, maxUnavailable: 50% was configured on a 4-pod cluster, instantly slashing available serving capacity in half during rollouts. Second, their container images lacked readinessProbe definitions; Kubernetes marked newly created pods as healthy the millisecond the container process launched, routing production traffic to pods that required 15 seconds to initialize database connections.

The platform team updated all production Deployment manifests: maxUnavailable was set strictly to 0, maxSurge was set to 25%, and rigorous HTTP readinessProbe endpoints were implemented. During the subsequent release, zero dropped connections and zero 502 errors occurred, achieving true zero-downtime rolling updates.

6. Common Pitfalls, Anti-Patterns & Misconceptions

Avoid these widespread Kubernetes workload anti-patterns:

  • Omitting Resource Requests and Limits: Running Pods without resource requests disables intelligent scheduling, allowing noisy neighbors to exhaust node memory and trigger kernel OOM kills. Remedy: Always specify both requests and limits for CPU and memory.
  • Confusing Liveness and Readiness Probes: Using a liveness probe to check downstream external dependencies (like an external database) causes all Pods to fail and restart simultaneously if the database encounters momentary latency. Remedy: Liveness probes should only verify local process responsiveness; readiness probes handle traffic routability.
  • Using the :latest Image Tag: Deploying containers with image: my-app:latest breaks deterministic deployments and impedes automated rollbacks. Remedy: Pin images to immutable semantic version tags or git commit SHAs.
  • Storing State Inside Pod Filesystems: Writing files to the ephemeral pod container filesystem causes permanent data loss when pods reschedule or restart. Remedy: Attach dedicated PersistentVolumeClaims for stateful persistence.

7. Best Practices, Security Hardening & Performance Checklists

Adhere to this enterprise production checklist for Kubernetes workloads:

  • Enforce Non-Root Execution: Configure securityContext.runAsNonRoot: true and drop unnecessary Linux capabilities (capabilities: { drop: ["ALL"] }) in the PodSpec.
  • Configure PodDisruptionBudgets (PDB): Define PDBs (e.g., minAvailable: 2) to ensure voluntary cluster disruptions (like node draining during node upgrades) do not violate service availability agreements.
  • Pod Anti-Affinity: Apply podAntiAffinity rules using topologyKey: "kubernetes.io/hostname" to guarantee pod replicas are distributed across distinct physical worker nodes and availability zones.
  • Graceful Shutdown Termination: Handle the SIGTERM signal in your application, delaying process exit by 5–10 seconds to allow the ingress controller to remove the Pod’s endpoint from active routing tables.

8. Summary & Certification Readiness Review

The SkillCertify Kubernetes Fundamentals Credential assessment tests candidates on control plane component roles (kube-apiserver, etcd, scheduler, controller-manager), multi-container pod design, rolling update deployment parameters (maxSurge and maxUnavailable), and health probe mechanics. Candidates must be prepared to troubleshoot broken manifests and rollout failures. Study the official Kubernetes documentation resources below to prepare thoroughly for your assessment.

Formative Practice

Test Your Understanding of Core K8s Workloads

Apply what you just learned with curated practice questions and in-depth explanations.

Practice Questions →
Advertisement