Nomad集成Azure Disk CSI插件报错:无法找到节点
Nomad v1.9.4中Azure Disk CSI节点插件健康状态异常排查
问题概述
在Nomad v1.9.4环境下,使用Azure托管身份对接Azure Disk CSI插件,控制器和节点任务已成功启动,但节点任务始终无法进入健康状态。日志显示插件查询CSI时使用了错误的节点名称,且尝试与Kubernetes交互导致报错。
错误日志
I0423 01:28:41.199953 1 utils.go:105] GRPC call: /csi.v1.Identity/GetPluginCapabilities I0423 01:28:41.199977 1 utils.go:106] GRPC request: {} I0423 01:28:41.200008 1 utils.go:112] GRPC response: {"capabilities":[{"Type":{"Service":{"type":1}}},{"Type":{"Service":{"type":2}}},{"Type":{"VolumeExpansion":{"type":2}}},{"Type":{"VolumeExpansion":{"type":1}}]} I0423 01:28:41.200521 1 utils.go:105] GRPC call: /csi.v1.Node/NodeGetInfo I0423 01:28:41.200532 1 utils.go:106] GRPC request: {} I0423 01:28:41.200560 1 azure_vmss_cache.go:425] Node 11c058f08c97 has joined the cluster since the last VM cache refresh in NonVmssUniformNodesEntry, refreshing the cache I0423 01:28:41.200568 1 azure_vmss_cache.go:342] refresh the cache of NonVmssUniformNodesCache in rg &{map[<redacted_correct_node_name>:{}]} I0423 01:28:41.251795 1 azure_vmss.go:811] Couldn't find VMSS for node 11c058f08c97, refreshing the cache W0423 01:28:41.317305 1 azure_vmss.go:817] Unable to find node 11c058f08c97: instance not found W0423 01:28:41.317330 1 nodeserver.go:359] get zone(<redacted_correct_node_name>) failed with: instance not found, fall back to get zone from node labels E0423 01:28:41.317343 1 utils.go:110] GRPC error: rpc error: code = Internal desc = getNodeInfoFromLabels on node(<redacted_correct_node_name>) failed with kubeClient is nil I0423 01:29:11.319481 1 utils.go:105] GRPC call: /csi.v1.Identity/GetPluginCapabilities I0423 01:29:11.319526 1 utils.go:106] GRPC request: {} I0423 01:29:11.319576 1 utils.go:112] GRPC response: {"capabilities":[{"Type":{"Service":{"type":1}}},{"Type":{"Service":{"type":2}}},{"Type":{"VolumeExpansion":{"type":2}}},{"Type":{"VolumeExpansion":{"type":1}}]} I0423 01:29:11.320740 1 utils.go:105] GRPC call: /csi.v1.Node/NodeGetInfo I0423 01:29:11.320757 1 utils.go:106] GRPC request: {} I0423 01:29:11.320971 1 azure_vmss_cache.go:425] Node 11c058f08c97 has joined the cluster since the last VM cache refresh in NonVmssUniformNodesEntry, refreshing the cache I0423 01:29:11.321131 1 azure_vmss_cache.go:342] refresh the cache of NonVmssUniformNodesCache in rg &{map[<redacted_correct_node_resource_group>:{}]} I0423 01:29:11.409074 1 azure_vmss.go:811] Couldn't find VMSS for node 11c058f08c97, refreshing the cache W0423 01:29:11.473572 1 azure_vmss.go:817] Unable to find node 11c058f08c97: instance not found W0423 01:29:11.473594 1 nodeserver.go:359] get zone(<redacted_correct_node_name>) failed with: instance not found, fall back to get zone from node labels E0423 01:29:11.473607 1 utils.go:110] GRPC error: rpc error: code = Internal desc = getNodeInfoFromLabels on node(<redacted_correct_node_name>) failed with kubeClient is nil
节点任务定义
job "plugin-azure-disk-nodes" { datacenters = ["dc1"] # you can run node plugins as service jobs as well, but this ensures # that all nodes in the DC have a copy. type = "system" constraint { attribute = "${node.unique.name}" operator = "regexp" value = "${node_regex}" } group "nodes" { task "node" { driver = "docker" template { change_mode = "noop" destination = "local/azure.json" data = <<EOH { "cloud":"AzurePublicCloud", "location": "${azure_resource_group_location}", "resourceGroup": "${azure_resource_group_name}", "subscriptionId": "${azure_subscription_id}", "tenantId": "${azure_tenant_id}", "useManagedIdentityExtension": true } EOH } env { AZURE_CREDENTIAL_FILE = "/etc/kubernetes/azure.json" } config { image = "mcr.microsoft.com/oss/kubernetes-csi/azuredisk-csi:v1.32.1" volumes = [ "local/azure.json:/etc/kubernetes/azure.json" ] args = [ "--nodeid=${node.unique.name}", "--endpoint=unix://csi/csi.sock", "--logtostderr", "--v=5", ] # node plugins must run as privileged jobs because they # mount disks to the host privileged = true } csi_plugin { id = "azure-disk" type = "node" mount_dir = "/csi" } resources { memory = 256 } # ensuring the plugin has time to shut down gracefully kill_timeout = "2m" } } }
解决方案建议
修正NodeID参数:Azure Disk CSI插件需要的是Azure VM的实例名称,而非Nomad自动生成的
node.unique.name。可通过Nomad节点属性传递实例名(如${node.attr.azure.instance-name},需提前配置节点属性),或通过模板从Azure元数据服务获取:template { change_mode = "noop" destination = "local/node-id" data = <<EOH {{- with get "http://169.254.169.254/metadata/instance/compute/name?api-version=2021-02-01&format=text" -}} {{ .Content | trimSpace }} {{- end -}} EOH }然后在args中使用
--nodeid=@local/node-id读取该文件内容。禁用Kubernetes相关逻辑:添加
--kubeconfig=""到插件args中,或设置环境变量KUBECONFIG="",阻止插件尝试初始化Kubernetes客户端。部分版本还可添加--disable-kubernetes-watch=true参数彻底关闭K8s相关监听。验证托管身份权限:确保Nomad节点的托管身份拥有以下权限:
- 资源组范围内的
Microsoft.Compute/disks/*权限 Microsoft.Compute/virtualMachines/read权限(用于查询VM实例信息)
- 资源组范围内的
检查元数据服务访问:确认Nomad节点可访问Azure元数据服务
http://169.254.169.254,无防火墙规则或代理拦截该请求。
内容的提问来源于stack exchange,提问作者Joe
相关产品推荐
相关产品推荐

