Terraform部署K8s应用失败:副本无法就绪问题求助
Terraform部署K8s资源失败问题排查与修复
问题现象
部署Terraform项目时触发错误:
Error: Waiting for rollout to finish: 3 replicas wanted; 0 replicas Ready
已配置RollingUpdate策略但问题未解决,kubectl查询结果:
- app1 Pod:部分Pending、部分ImagePullBackOff
- app2 Pod:部分Pending、部分ImagePullBackOff
- app3 Pod:全部InvalidImageName
- 所有Deployment均超过进度截止时间
相关配置代码
部署配置
resource "kubernetes_deployment" "dep_apps" { for_each = var.apps metadata { name = each.value.appName namespace = each.value.appName labels = { name = each.value.labels.name tier = each.value.labels.tier } } spec { replicas = 3 strategy { type = "RollingUpdate" rolling_update { max_surge = "25%" max_unavailable = "25%" } } selector { match_labels = { name = each.value.labels.name tier = each.value.labels.tier } } template { metadata { name = each.value.appName namespace = each.value.appName labels = { name = each.value.labels.name tier = each.value.labels.tier } } spec { container { name = each.value.appName image = each.value.image resources { limits = { cpu = "500m" memory = "512Mi" } requests = { cpu = "200m" memory = "256Mi" } } } } } } }
HPA配置
resource "kubernetes_horizontal_pod_autoscaler_v1" "autoscaler" { for_each = var.apps metadata { name = "${each.value.appName}-as" namespace = each.value.appName labels = { name = each.value.labels.name tier = each.value.labels.tier } } spec { scale_target_ref { api_version = "apps/v1" kind = "ReplicaSet" name = each.value.appName } min_replicas = 1 max_replicas = 10 } }
Service配置
resource "kubernetes_service" "load_balancer" { for_each = toset([for app in var.apps : app if app.appName == ["app1", "app2"]]) metadata { name = "${each.value.appName}-lb" labels = { name = each.value.labels.name tier = each.value.labels.tier } } spec { selector = { app = each.value.appName } port { name = "http" port = 80 target_port = 8080 } type = "LoadBalancer" } }
变量定义
variable "apps" { type = map(object({ appName = string team = string labels = map(string) annotations = map(string) data = map(string) image = string })) default = { "app1" = { appName = "app1" team = "frontend" image = "nxinx" labels = { "name" = "stream-frontend" "tier" = "web" "owner" = "product" } annotations = { "serviceClass" = "web-frontend" "loadBalancer_and_class" = "external" } data = { "aclName" = "acl_frontend" "ingress" = "stream-frontend" "egress" = "0.0.0.0/0" "port" = "8080" "protocol" = "TCP" } } "app2" = { appName = "app2" team = "backend" image = "nginx:dev" labels = { "name" = "stream-frontend" "tier" = "web" "owner" = "product" } annotations = { "serviceClass" = "web-frontend" "loadBalancer_and_class" = "external" } data = { "aclName" = "acl_backend" "ingress" = "stream-backend" "egress" = "0.0.0.0/0" "port" = "8080" "protocol" = "TCP" } } "app3" = { appName = "app3" team = "database" image = "Mongo" labels = { "name" = "stream-database" "tier" = "shared" "owner" = "product" } annotations = { "serviceClass" = "disabled" "loadBalancer_and_class" = "disabled" } data = { "aclName" = "acl_database" "ingress" = "stream-database" "egress" = "172.17.0.0/24" "port" = "27017" "protocol" = "TCP" } } } }
kubectl查询结果
kubectl get deployments NAME READY UP-TO-DATE AVAILABLE AGE app1 0/3 3 0 34m kubectl rollout status deployment app1 error: deployment "app1" exceeded its progress deadline kubectl get namespace NAME STATUS AGE app1 Active 6h16m app2 Active 6h16m app3 Active 6h16m default Active 4d15h kube-node-lease Active 4d15h kube-public Active 4d15h kube-system Active 4d15h kubectl get pods NAME READY STATUS RESTARTS AGE app1-7f8657489c-579lm 0/1 Pending 0 37m app1-7f8657489c-jjppq 0/1 ImagePullBackOff 0 37m app1-7f8657489c-lv49l 0/1 ImagePullBackOff 0 37m kubectl get pods --namespace=app2 NAME READY STATUS RESTARTS AGE app2-68b6b59584-86dt8 0/1 ImagePullBackOff 0 38m app2-68b6b59584-8kr2p 0/1 Pending 0 38m app2-68b6b59584-jzzxt 0/1 Pending 0 38m kubectl get pods --namespace=app3 NAME READY STATUS RESTARTS AGE app3-5f589dc88d-gwn2n 0/1 InvalidImageName 0 39m app3-5f589dc88d-pzhzw 0/1 InvalidImageName 0 39m app3-5f589dc88d-vx452 0/1 InvalidImageName 0 39m
修复方案
1. 修正镜像名称
这是Pod启动失败的核心原因:
- app1镜像改为
nginx(原拼写错误为nxinx) - app2镜像改为
nginx:latest(nginx:dev标签不存在) - app3镜像改为
mongo(Docker Hub镜像名需小写,原Mongo不符合规范)
修改后的变量中镜像字段:
"app1" = { # ... image = "nginx" # ... } "app2" = { # ... image = "nginx:latest" # ... } "app3" = { # ... image = "mongo" # ... }
2. 修复HPA配置
HPA需指向Deployment而非ReplicaSet,修正scale_target_ref的kind:
resource "kubernetes_horizontal_pod_autoscaler_v1" "autoscaler" { # ... spec { scale_target_ref { api_version = "apps/v1" kind = "Deployment" # 从ReplicaSet改为Deployment name = each.value.appName } # ... } }
3. 修正Service选择器
Service的selector需与Pod的标签匹配,原配置用了不存在的app标签,改为Pod实际的name和tier:
resource "kubernetes_service" "load_balancer" { # ... spec { selector = { name = each.value.labels.name tier = each.value.labels.tier } # ... } }
4. 修复Service的for_each逻辑
原判断条件app.appName == ["app1", "app2"]逻辑错误,改用contains函数筛选目标应用:
resource "kubernetes_service" "load_balancer" { for_each = { for k, app in var.apps : k => app if contains(["app1", "app2"], app.appName) } # ... }
5. 排查Pod Pending状态
执行以下命令查看Pending Pod的详细事件,确认是资源不足、存储卷配置还是调度问题:
kubectl describe pod <pod-name> --namespace=<namespace>
根据事件提示调整节点资源分配或修复相关配置。
内容的提问来源于stack exchange,提问作者zolo
相关产品推荐
相关产品推荐

