⚡ DevOps Arcade
☸️ Kubernetes SRE Incident Drill

kubectl Crash Simulator

Production is failing. Alerts are firing. Diagnose and fix CrashLoopBackOff, Exit Code 137 OOMKills, Pending PVCs, and broken Services under a 60-second timer.

Incident #1 of 10
Time Remaining 60s
SRE Score 0 PTS
Streak 0x 🔥
Triage Rank L1 Operator
bash — naveed@k8s-prod-cluster-01 — 112x32
ns: prod
Step 1: Choose Diagnostic Command What is your first triage command?
$
Quick Diagnostic Commands:
kubectl get pods -n prod -o wide kubectl describe pod kubectl logs --previous kubectl get endpoints kubectl get events
⭐ Staff SRE (Level 4/5)

Cluster Outage Triage Complete

You diagnosed and resolved 8 of 10 production cluster incidents under live timer conditions.

Final Score
3,850
Incidents Fixed
8 / 10
Mean Time to Detect (MTTD)
24.2s
Max Streak
5x 🔥

The SRE 4-Step Algorithmic Kubernetes Triage Framework

When an outage alerts you at 2:00 AM, guessing and deleting pods wastes critical time. Follow this deterministic triage sequence engineered for production reliability.

[ALERT: Pod Not Ready / AlertManager Fired] │ ▼ [Step 1: Scope the Blast Radius] $ kubectl get pods -n <namespace> -o wide │ ├──► Is only 1 Pod crashing? ──────────────► Application Config / Env Var / Code Bug └──► Are all Pods on 1 Node failing? ──────► Worker Node / Kubelet / CNI Network Down │ ▼ [Step 2: Inspect Container Life Cycle] $ kubectl describe pod <pod-name> -n <namespace> │ ├──► Check "Last State: Terminated" (Look for Exit Code & Reason) ├──► Check "Events" at bottom (FailedScheduling, BackOff, Unhealthy) └──► Verify "Limits & Requests" (Is memory or CPU saturated?) │ ▼ [Step 3: Capture Post-Mortem Crash Logs] $ kubectl logs <pod-name> -n <namespace> --previous --tail=100 │ └──► Read fatal stack trace right before the process terminated │ ▼ [Step 4: Verify Network Endpoints & DNS] $ kubectl get endpoints <service-name> -n <namespace> $ kubectl get events -n <namespace> --sort-by='.metadata.creationTimestamp'

Exit Code Rosetta Stone: Decoding Kernel & Kubelet Signals

Before inspecting application logs, the container exit code provides immediate insight into the root failure vector.

Exit Code Signal Name Primary Root Cause First Diagnostic Command
137 SIGKILL (128 + 9) OOMKilled: Container memory exceeded cgroup limits or node ran out of RAM. kubectl describe pod <pod> | grep -i oom
143 SIGTERM (128 + 15) Graceful Kill Interrupted: Failed liveness probe or pod eviction timeout. Check liveness probe initialDelaySeconds and grace periods.
1 Application Error Runtime Panic: Missing required env var, DB handshake error, or syntax bug. kubectl logs <pod> --previous
127 Command Not Found Missing Binary: Entrypoint or command executable not found in image $PATH. Review Dockerfile base image and script shebang.
139 SIGSEGV (128 + 11) Segmentation Fault: Memory corruption or C-library mismatch (Alpine vs glibc). Inspect native compiled dependencies and glibc versions.
0 Success / Clean Exit Premature Completion: Daemon launched in background; main process exited. Ensure foreground execution flag (e.g. nginx -g 'daemon off;').

📚 Frequently Asked Kubernetes Triage Questions

What is the difference between CrashLoopBackOff and an application error? +
CrashLoopBackOff is not an error code; it is a Kubernetes state condition. When a container process terminates unexpectedly with a non-zero exit code, the kubelet attempts to restart it. If it crashes repeatedly, the kubelet applies an exponential back-off delay (10s, 20s, 40s up to 5 minutes) before the next restart to avoid overloading node CPU and disk. The application error is the underlying cause, while CrashLoopBackOff is the orchestrator's pacing response.
How do you diagnose an Exit Code 137 in Kubernetes? +
Exit Code 137 indicates that the Linux kernel sent a SIGKILL signal (128 + 9 = 137) to terminate the container. The primary cause is OOMKilled (Out Of Memory), which happens when a container exceeds the memory limit defined in resources.limits.memory. Run kubectl describe pod <pod-name> and inspect the Last State: Terminated block. If Reason: OOMKilled is present, the process breached its cgroup limit. Resolution requires profiling memory leaks or increasing the container memory limit.
Why does kubectl logs show no output when a pod is in CrashLoopBackOff? +
Running kubectl logs <pod-name> queries the currently running container instance. If the container is currently waiting in its back-off period between restarts, the active container does not exist or has not yet written logs. To inspect the stdout and stderr logs from the failed container instance right before it crashed, pass the --previous flag: kubectl logs <pod-name> --previous.
How do you triage a Kubernetes Service returning HTTP 503 errors when pods are Running? +
If pods show 1/1 Running but the Service returns HTTP 503, the issue is almost always a label selector mismatch. Run kubectl get endpoints <service-name> or kubectl get endpointslices. If the ENDPOINTS column shows <none>, the Service's spec.selector labels do not match the Pod's metadata.labels. Align the labels so kube-proxy can register the pod IPs as healthy backend endpoints.
What causes a Kubernetes Pod to stay in Pending state indefinitely? +
A Pod stays in Pending when the Kubernetes scheduler cannot place it on any worker node. Run kubectl describe pod <pod-name> and inspect the Events section at the bottom. Common causes include: 1) Insufficient CPU or Memory requests exceeding available allocatable capacity on all nodes; 2) PersistentVolumeClaim (PVC) unbound due to storageClass mismatch or volume binding mode; 3) Unmatched nodeSelector, nodeAffinity, or taints without matching tolerations.
What is the difference between Exit Code 143 and Exit Code 137? +
Exit Code 143 corresponds to SIGTERM (128 + 15 = 143), which is a graceful termination signal sent by the kubelet (e.g. during pod eviction, rolling update, or when a liveness probe fails). The application is given a grace period (default 30s) to shut down cleanly. Exit Code 137 corresponds to SIGKILL (128 + 9 = 137), an immediate, ungraceful kill triggered by the Linux OOM killer when cgroup memory limits are exceeded.