← ClaudeAtlas

diag-k8s-pod-crashlooplisted

K8s Pod CrashLoopBackOff 全链排查。从 kubectl describe → events → logs → restart policy → resource limits → liveness/readiness probe → image pull → configmap/secret 挂载,���位根因并输出修复建议。
seed-forge/harness-ai-kit · ★ 22 · AI & Automation · score 74
Install: claude install-skill seed-forge/harness-ai-kit
# K8s Pod CrashLoopBackOff Full-Chain Diagnostics ## 用途 当 Pod 反复重启、状态显示 `CrashLoopBackOff` 或 `Error` 时触发。本技能是**自包含诊断 Runbook**。 ## 输入 - kubectl 上下文(kubeconfig) - 目标 Pod 名称 + Namespace ## 输出 - CrashLoop 根因分析报告 + 修复建议 ## 诊断步骤 ### Step 1: Pod 状态与事件 ```bash kubectl get pod {pod_name} -n {namespace} -o wide kubectl describe pod {pod_name} -n {namespace} # 关注 Events 区块:OOMKilled / Error / ImagePullBackOff / FailedMount ``` ### Step 2: 容器日志 ```bash # 当前容器日志(可能在 crash 后已清空) kubectl logs {pod_name} -n {namespace} --all-containers=true --tail=100 # 上一次崩溃的日志 kubectl logs {pod_name} -n {namespace} --all-containers=true --previous --tail=200 ``` ### Step 3: 重启原因分类 | 事件关键词 | 根因 | 典型修复 | |-----------|------|---------| | `OOMKilled` | 内存超限 | 增大 `resources.limits.memory` 或修复内存泄漏 | | `Error` (exit code 1) | 应用启动失败 | 检查日志、配置、依赖服务连通性 | | `ImagePullBackOff` | 镜像拉取失败 | 检查 image tag、imagePullSecrets、registry 可达 | | `FailedMount` | ConfigMap/Secret/PVC 挂载失败 | 检查引用是否存在 | | `Liveness probe failed` | 健康检查超时 | 调整 `initialDelaySeconds` / `timeoutSeconds` | | `CreateContainerConfigError` | env/configMapKeyRef 缺失 | 检查引用的 ConfigMap/Secret key | | `BackOff restarting` | 重启间隔递增 | 以上任一原因的持续重试 | ### Step 4: 资源限制与请求 ```bash kubectl get pod {pod_name} -n {namespace} -o jsonpath='{range .spec.containers[*]}{.name}: requests={.resources.requests} limits={.resources.limits}{"\n"}{end}' ``` ### Step 5: 依赖检查 ```bash # ConfigMap/Secret 是否存在 kubectl get configmap -n {namespace} | grep {configmap_name}