You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Nomad集成Azure Disk CSI插件报错:无法找到节点

Nomad v1.9.4中Azure Disk CSI节点插件健康状态异常排查

问题概述

在Nomad v1.9.4环境下,使用Azure托管身份对接Azure Disk CSI插件,控制器和节点任务已成功启动,但节点任务始终无法进入健康状态。日志显示插件查询CSI时使用了错误的节点名称,且尝试与Kubernetes交互导致报错。

错误日志

I0423 01:28:41.199953       1 utils.go:105] GRPC call: /csi.v1.Identity/GetPluginCapabilities
I0423 01:28:41.199977       1 utils.go:106] GRPC request: {}
I0423 01:28:41.200008       1 utils.go:112] GRPC response: {"capabilities":[{"Type":{"Service":{"type":1}}},{"Type":{"Service":{"type":2}}},{"Type":{"VolumeExpansion":{"type":2}}},{"Type":{"VolumeExpansion":{"type":1}}]} 
I0423 01:28:41.200521       1 utils.go:105] GRPC call: /csi.v1.Node/NodeGetInfo
I0423 01:28:41.200532       1 utils.go:106] GRPC request: {}
I0423 01:28:41.200560       1 azure_vmss_cache.go:425] Node 11c058f08c97 has joined the cluster since the last VM cache refresh in NonVmssUniformNodesEntry, refreshing the cache
I0423 01:28:41.200568       1 azure_vmss_cache.go:342] refresh the cache of NonVmssUniformNodesCache in rg &{map[<redacted_correct_node_name>:{}]}
I0423 01:28:41.251795       1 azure_vmss.go:811] Couldn't find VMSS for node 11c058f08c97, refreshing the cache
W0423 01:28:41.317305       1 azure_vmss.go:817] Unable to find node 11c058f08c97: instance not found
W0423 01:28:41.317330       1 nodeserver.go:359] get zone(<redacted_correct_node_name>) failed with: instance not found, fall back to get zone from node labels
E0423 01:28:41.317343       1 utils.go:110] GRPC error: rpc error: code = Internal desc = getNodeInfoFromLabels on node(<redacted_correct_node_name>) failed with kubeClient is nil
I0423 01:29:11.319481       1 utils.go:105] GRPC call: /csi.v1.Identity/GetPluginCapabilities
I0423 01:29:11.319526       1 utils.go:106] GRPC request: {}
I0423 01:29:11.319576       1 utils.go:112] GRPC response: {"capabilities":[{"Type":{"Service":{"type":1}}},{"Type":{"Service":{"type":2}}},{"Type":{"VolumeExpansion":{"type":2}}},{"Type":{"VolumeExpansion":{"type":1}}]} 
I0423 01:29:11.320740       1 utils.go:105] GRPC call: /csi.v1.Node/NodeGetInfo
I0423 01:29:11.320757       1 utils.go:106] GRPC request: {}
I0423 01:29:11.320971       1 azure_vmss_cache.go:425] Node 11c058f08c97 has joined the cluster since the last VM cache refresh in NonVmssUniformNodesEntry, refreshing the cache
I0423 01:29:11.321131       1 azure_vmss_cache.go:342] refresh the cache of NonVmssUniformNodesCache in rg &{map[<redacted_correct_node_resource_group>:{}]}
I0423 01:29:11.409074       1 azure_vmss.go:811] Couldn't find VMSS for node 11c058f08c97, refreshing the cache
W0423 01:29:11.473572       1 azure_vmss.go:817] Unable to find node 11c058f08c97: instance not found
W0423 01:29:11.473594       1 nodeserver.go:359] get zone(<redacted_correct_node_name>) failed with: instance not found, fall back to get zone from node labels
E0423 01:29:11.473607       1 utils.go:110] GRPC error: rpc error: code = Internal desc = getNodeInfoFromLabels on node(<redacted_correct_node_name>) failed with kubeClient is nil

节点任务定义

job "plugin-azure-disk-nodes" {
  datacenters = ["dc1"]

  # you can run node plugins as service jobs as well, but this ensures
  # that all nodes in the DC have a copy.
  type = "system"

  constraint {
    attribute = "${node.unique.name}"
    operator  = "regexp"
    value     = "${node_regex}"
  }

  group "nodes" {
    task "node" {
      driver = "docker"

      template {
        change_mode = "noop"
        destination = "local/azure.json"
        data = <<EOH
{
"cloud":"AzurePublicCloud",
"location": "${azure_resource_group_location}",
"resourceGroup": "${azure_resource_group_name}",
"subscriptionId": "${azure_subscription_id}",
"tenantId": "${azure_tenant_id}",
"useManagedIdentityExtension": true
}
EOH
      }

      env {
        AZURE_CREDENTIAL_FILE = "/etc/kubernetes/azure.json"
      }

      config {
        image = "mcr.microsoft.com/oss/kubernetes-csi/azuredisk-csi:v1.32.1"

        volumes = [
          "local/azure.json:/etc/kubernetes/azure.json"
        ]

        args = [
          "--nodeid=${node.unique.name}",
          "--endpoint=unix://csi/csi.sock",
          "--logtostderr",
          "--v=5",
        ]

        # node plugins must run as privileged jobs because they
        # mount disks to the host
        privileged = true
      }

      csi_plugin {
        id        = "azure-disk"
        type      = "node"
        mount_dir = "/csi"
      }

      resources {
        memory = 256
      }

      # ensuring the plugin has time to shut down gracefully
      kill_timeout = "2m"
    }
  }
}

解决方案建议

  • 修正NodeID参数:Azure Disk CSI插件需要的是Azure VM的实例名称,而非Nomad自动生成的node.unique.name。可通过Nomad节点属性传递实例名(如${node.attr.azure.instance-name},需提前配置节点属性),或通过模板从Azure元数据服务获取:

    template {
      change_mode = "noop"
      destination = "local/node-id"
      data = <<EOH
    {{- with get "http://169.254.169.254/metadata/instance/compute/name?api-version=2021-02-01&format=text" -}}
    {{ .Content | trimSpace }}
    {{- end -}}
    EOH
    }
    

    然后在args中使用--nodeid=@local/node-id读取该文件内容。

  • 禁用Kubernetes相关逻辑:添加--kubeconfig=""到插件args中,或设置环境变量KUBECONFIG="",阻止插件尝试初始化Kubernetes客户端。部分版本还可添加--disable-kubernetes-watch=true参数彻底关闭K8s相关监听。

  • 验证托管身份权限:确保Nomad节点的托管身份拥有以下权限:

    • 资源组范围内的Microsoft.Compute/disks/*权限
    • Microsoft.Compute/virtualMachines/read权限(用于查询VM实例信息)
  • 检查元数据服务访问:确认Nomad节点可访问Azure元数据服务http://169.254.169.254,无防火墙规则或代理拦截该请求。

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 07:32:07