K3s中Terraform部署PVC与Pod陷入调度循环问题求助
问题描述
通过Terraform在K3s(使用local-path存储类)部署Grafana和Prometheus时,出现PVC与Pod创建的无限循环:
grafana-pvc处于WaitForFirstConsumer等待状态,需Pod调度后才能绑定动态生成的PV- 关联的Grafana Pod因找不到
grafana-configurationPVC而调度失败 - 最终Terraform超时,无法完成部署或恢复状态
grafana-pvc状态
Name: grafana-pvc Namespace: default StorageClass: local-path Status: Pending Volume: Labels: io.kompose.service=grafana-data Annotations: <none> Finalizers: [kubernetes.io/pvc-protection] Capacity: Access Modes: VolumeMode: Filesystem Used By: grafana-778c7f77c7-w7x9f Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal WaitForFirstConsumer 79s persistentvolume-controller waiting for first consumer to be created before binding Normal WaitForPodScheduled 7s (x5 over 67s) persistentvolume-controller waiting for pod grafana-778c7f77c7-w7x9f to be scheduled
Grafana Pod状态
Name: grafana-778c7f77c7-w7x9f Namespace: default Priority: 0 Service Account: default Node: <none> Labels: io.kompose.service=grafana pod-template-hash=778c7f77c7 Annotations: <none> Status: Pending IP: IPs: <none> Controlled By: ReplicaSet/grafana-778c7f77c7 Containers: grafana: Image: grafana/grafana:9.2.4 Port: 3000/TCP Host Port: 0/TCP Environment: <none> Mounts: /etc/grafana from grafana-configuration (rw) /var/lib/grafana from grafana-data (rw) /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-n7cmt (ro) Conditions: Type Status PodScheduled False Volumes: grafana-configuration: Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace) ClaimName: grafana-configuration ReadOnly: false grafana-data: Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace) ClaimName: grafana-pvc ReadOnly: false kube-api-access-n7cmt: Type: Projected (a volume that contains injected data from multiple sources) TokenExpirationSeconds: 3607 ConfigMapName: kube-root-ca.crt ConfigMapOptional: <nil> DownwardAPI: true QoS Class: BestEffort Node-Selectors: <none> Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s node.kubernetes.io/unreachable:NoExecute op=Exists for 300s Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning FailedScheduling 90s default-scheduler 0/1 nodes are available: 1 persistentvolumeclaim "grafana-configuration" not found. preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling. Warning FailedScheduling 89s default-scheduler 0/1 nodes are available: 1 persistentvolumeclaim "grafana-configuration" not found. preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling.
Terraform配置文件
grafana.tf
resource "kubernetes_persistent_volume_claim" "grafana-configuration" { metadata { name = "grafana-configuration" labels = { "io.kompose.service" = "grafana-configuration" } } spec { access_modes = ["ReadWriteOnce"] storage_class_name = "local-path" resources { requests = { storage = "1Gi" } } volume_name = "grafana-configuration" } } resource "kubernetes_persistent_volume" "grafana-configuration" { metadata { name = "grafana-configuration" } spec { storage_class_name = "local-path" access_modes = ["ReadWriteOnce"] capacity = { storage = "1Gi" } node_affinity { required { node_selector_term { match_expressions { key = "node-role.kubernetes.io/master" operator = "In" values = ["true"] } } } } persistent_volume_source { local { path = "/home/administrator/Metrics.Infrastructure/grafana/" } } } } resource "kubernetes_persistent_volume_claim" "grafana-pvc" { metadata { name = "grafana-pvc" labels = { "io.kompose.service" = "grafana-data" } } spec { access_modes = ["ReadWriteOnce"] storage_class_name = "local-path" resources { requests = { storage = "5Gi" } } } } resource "kubernetes_deployment" "grafana" { metadata { name = "grafana" labels = { "io.kompose.service" = "grafana" } } spec { replicas = 1 selector { match_labels = { "io.kompose.service" = "grafana" } } template { metadata { labels = { "io.kompose.service" = "grafana" } } spec { volume { name = "grafana-configuration" persistent_volume_claim { claim_name = "grafana-configuration" } } volume { name = "grafana-data" persistent_volume_claim { claim_name = "grafana-pvc" } } container { name = "grafana" image = "grafana/grafana:9.2.4" port { container_port = 3000 } volume_mount { name = "grafana-configuration" mount_path = "/etc/grafana" } volume_mount { name = "grafana-data" mount_path = "/var/lib/grafana" } } restart_policy = "Always" } } strategy { type = "Recreate" } } } resource "kubernetes_service" "grafana" { metadata { name = "grafana" labels = { "io.kompose.service" = "grafana" } } spec { port { port = 3000 target_port = 3000 node_port = 30001 } type = "NodePort" selector = { "io.kompose.service" = "grafana" } } }
prometheus.tf
# We need these resources so that prometheus can fetch kubernetes metrics resource "kubernetes_cluster_role" "prometheus-clusterrole" { metadata { name = "prometheus-clusterrole" } rule { api_groups = [""] resources = ["nodes", "nodes/proxy", "services", "endpoints", "pods"] verbs = ["get", "list", "watch"] } rule { api_groups = ["extensions"] resources = ["ingresses"] verbs = ["get", "list", "watch"] } rule { non_resource_urls = ["/metrics"] verbs = ["get"] } } resource "kubernetes_cluster_role_binding" "prometheus_clusterrolebinding" { metadata { name = "prometheus-clusterrolebinding" } role_ref { api_group = "rbac.authorization.k8s.io" kind = "ClusterRole" name = "prometheus-clusterrole" } subject { kind = "ServiceAccount" name = "default" namespace = "default" } } resource "kubernetes_config_map" "prometheus-config" { metadata { name = "prometheus-config" } data = { "prometheus.yml" = "${file("${path.module}/prometheus/prometheus.yml")}" } } resource "kubernetes_persistent_volume_claim" "prometheus_data_claim" { metadata { name = "prometheus-data-claim" labels = { "io.kompose.service" = "prometheus-data" } } spec { access_modes = ["ReadWriteOnce"] storage_class_name = "local-path" resources { requests = { storage = "20Gi" } } } } resource "kubernetes_deployment" "prometheus" { metadata { name = "prometheus" labels = { "io.kompose.service" = "prometheus" } } spec { replicas = 1 selector { match_labels = { "io.kompose.service" = "prometheus" } } template { metadata { labels = { "io.kompose.service" = "prometheus" } } spec { volume { name = "prometheus-data" persistent_volume_claim { claim_name = "prometheus-data-claim" } } volume { name = "prometheus-config" config_map { name = "prometheus-config" } } container { name = "prometheus" image = "prom/prometheus:v2.40.0" args = [ "--config.file=/config/prometheus.yml", "--storage.tsdb.path=/prometheus", "--web.enable-lifecycle" ] port { container_port = 9090 } volume_mount { name = "prometheus-config" mount_path = "/config" } volume_mount { name = "prometheus-data" mount_path = "/prometheus" } } restart_policy = "Always" } } strategy { type = "Recreate" } } } resource "kubernetes_service" "prometheus" { metadata { name = "prometheus" labels = { "io.kompose.service" = "prometheus" } } spec { port { port = 80 target_port = 9090 node_port = 30000 } type = "NodePort" selector = { "io.kompose.service" = "prometheus" } } }
问题原因与修复方案
原因分析
- 资源依赖顺序错误:Terraform未明确Deployment与PVC的依赖关系,可能先创建Deployment,此时
grafana-configurationPVC尚未就绪,导致Pod调度失败。 - 手动PV与动态存储类冲突:手动创建
grafana-configurationPV并在PVC中指定volume_name,但local-path存储类默认是动态创建PV模式,手动绑定导致PVC无法正常就绪。 - 循环等待死锁:
grafana-pvc使用WaitForFirstConsumer模式,需Pod调度后才会生成PV;但Pod因缺失PVC无法调度,形成死循环。
修复步骤
1. 明确Deployment与PVC的依赖关系
在Grafana Deployment资源中添加depends_on,确保Terraform先创建PVC再部署Pod:
resource "kubernetes_deployment" "grafana" { # ... 现有配置 ... depends_on = [ kubernetes_persistent_volume_claim.grafana-configuration, kubernetes_persistent_volume_claim.grafana-pvc ] }
2. 修正grafana-configuration存储配置
如果不需要手动挂载节点路径,直接删除手动创建的PV,让local-path动态生成:
# 删除以下手动PV资源块 # resource "kubernetes_persistent_volume" "grafana-configuration" { # ... # }
同时修改对应的PVC,移除volume_name字段:
resource "kubernetes_persistent_volume_claim" "grafana-configuration" { metadata { name = "grafana-configuration" labels = { "io.kompose.service" = "grafana-configuration" } } spec { access_modes = ["ReadWriteOnce"] storage_class_name = "local-path" resources { requests = { storage = "1Gi" } } # 移除该行:volume_name = "grafana-configuration" } }
如果必须手动挂载节点路径,需确保:
- 节点上的
/home/administrator/Metrics.Infrastructure/grafana/路径存在,且权限适配Grafana容器默认用户(UID 472) - 手动PV的
storage_class_name与PVC保持一致,且PVC不依赖动态存储类
3. 重新部署验证
执行以下命令重新部署:
terraform plan terraform apply
部署完成后检查资源状态:
kubectl get pods,pvc
内容的提问来源于stack exchange,提问作者kaffarell
相关产品推荐
相关产品推荐

