Azure ML计算集群User Assigned Identity配置问题咨询
解决Azure ML Compute Cluster继承身份或使用当前用户身份的问题
一、让Compute Cluster使用指定的用户分配身份(与Compute Instance同身份)
Compute Cluster不会自动继承Compute Instance的身份,必须显式在创建时指定完整的用户分配身份ARM ID,而非仅实例化空的UserIdentityConfiguration()。
1. 获取用户分配身份的ARM ID
在你的Compute Instance中,通过以下命令获取目标身份的完整ID:
az identity list --query "[?name=='<你的用户分配身份名称>'].id" -o tsv
也可通过Azure SDK从当前Compute Instance的元数据中获取:
import requests metadata_url = "http://169.254.169.254/metadata/identity/info?api-version=2018-02-01" headers = {"Metadata": "true"} response = requests.get(metadata_url, headers=headers).json() identity_id = response['identityIds'][0] # 若实例绑定多个身份,需选择对应项
2. 使用Azure SDK v2正确配置集群身份
以Python SDK为例,创建集群时需将身份ID传入UserIdentityConfiguration的user_assigned_identities参数:
from azure.mgmt.machinelearningservices import MachineLearningServicesClient from azure.mgmt.machinelearningservices.models import AmlCompute, UserIdentityConfiguration, ComputeResource # 初始化ML客户端(已通过Compute Instance身份认证) ml_client = MachineLearningServicesClient.from_config() # 替换为你的资源信息和身份ID resource_group = "<资源组名称>" workspace_name = "<工作区名称>" cluster_name = "<集群名称>" identity_id = "<获取到的用户分配身份ARM ID>" # 构建集群配置 compute_config = AmlCompute( properties={ "vmSize": "Standard_DS3_v2", "scaleSettings": {"maxNodeCount": 4}, "identity": UserIdentityConfiguration( user_assigned_identities={identity_id: {}} ) } ) # 创建/更新集群 ml_client.compute.begin_create_or_update( resource_group_name=resource_group, workspace_name=workspace_name, compute_name=cluster_name, parameters=ComputeResource(properties=compute_config) ).result()
这样创建的集群,其yml配置中的identity字段会正确填充用户分配身份信息。
二、让Compute Cluster作业使用当前az login用户身份
如果希望集群运行作业时使用执行代码的az login用户身份,需启用身份传递(Identity Passthrough),无需配置集群自身的用户分配身份,而是在提交作业时指定用户身份:
1. 提交作业时配置用户身份
from azure.ai.ml import command, MLClient from azure.ai.ml.entities import UserIdentity # 初始化ML客户端(使用az login的用户身份认证) ml_client = MLClient.from_config() # 定义作业 job = command( code="./src", command="python train.py", environment="AzureML-sklearn-0.24-ubuntu18.04-py37-cpu@latest", compute="<集群名称>", identity=UserIdentity() # 指定使用提交作业的用户身份 ) # 提交作业 returned_job = ml_client.jobs.create_or_update(job)
作业运行时会自动使用提交者的身份访问Azure服务,前提是该用户拥有目标服务(CosmosDB、Databricks)的RBAC权限。
三、排查集群yml中identity为null的问题
- 检查SDK代码:确保没有仅实例化空的
UserIdentityConfiguration(),必须传入user_assigned_identities字典,键为身份ARM ID,值为空字典。 - 验证身份权限:确保Compute Instance使用的用户分配身份拥有以下权限:
Microsoft.MachineLearningServices/workspaces/computes/write(创建/更新集群)Microsoft.ManagedIdentity/userAssignedIdentities/assign/action(将身份分配给集群)
- 确认身份ID正确性:必须使用完整的ARM ID,格式为
/subscriptions/<订阅ID>/resourceGroups/<资源组>/providers/Microsoft.ManagedIdentity/userAssignedIdentities/<身份名称>
内容的提问来源于stack exchange,提问作者BeGreen
相关产品推荐
相关产品推荐

