Azure ML SDKv2:能否将注册数据集作为Choice传入调优作业?
问题解决:将注册数据集作为Choice传入Azure ML调优作业
可以实现将已注册的词嵌入数据集作为调优作业的Choice参数,错误原因是直接将Input对象或路径字符串放入Choice时,Azure ML会将其识别为字面量值,而非数据资产引用。以下是两种可行的实现方式:
方法1:使用数据集的ARM ID构造搜索空间
- 先获取每个已注册嵌入数据集的ARM ID:
from azure.ai.ml import MLClient from azure.identity import DefaultAzureCredential ml_client = MLClient(DefaultAzureCredential(), subscription_id="你的订阅ID", resource_group_name="资源组名", workspace_name="工作区名") # 获取每个数据集的ARM ID embsa = ml_client.data.get(name="embsa", version="1") embsb = ml_client.data.get(name="embsb", version="1") embsc = ml_client.data.get(name="embsc", version="1") embsd = ml_client.data.get(name="embsd", version="1") embs_arm_ids = [embsa.id, embsb.id, embsc.id, embsd.id]
- 在调优作业的搜索空间中,将ARM ID传入
Choice:
from azure.ai.ml.sweep import Choice # 定义搜索空间 search_space = { "initial_embedding": Choice(embs_arm_ids) }
- 确保你的微调组件的
initial_embedding输入类型定义为uri_file:
from azure.ai.ml import command # 定义微调组件 fine_tune_component = command( name="fine_tune_component", inputs={ "initial_embedding": Input(type="uri_file"), # 其他输入... }, command="python fine_tune.py --embedding ${{inputs.initial_embedding}}", # 其他组件配置... )
- 创建调优步骤并绑定参数:
from azure.ai.ml import sweep sweep_job = sweep( name="embedding_sweep", component=fine_tune_component, search_space=search_space, # 调优配置(比如采样方法、目标指标等) sampling_algorithm="random", objective={"primary_metric": "accuracy", "goal": "maximize"}, # 其他作业配置... )
方法2:直接使用资产引用字符串构造Choice
无需提前获取ARM ID,直接使用azureml:数据集名:版本格式的字符串传入Choice,Azure ML会自动解析该引用为对应的uri_file资产,只需确保组件输入类型为uri_file:
from azure.ai.ml.sweep import Choice search_space = { "initial_embedding": Choice([ "azureml:embsa:1", "azureml:embsb:1", "azureml:embsc:1", "azureml:embsd:1" ]) } # 后续组件定义和调优步骤与方法1一致
关键说明
- 不能直接将
Input对象放入Choice,搜索空间的参数值需要是可被Azure ML解析为资产的标识符(ARM ID或资产引用字符串),而非Input实例。 - 组件的输入必须明确指定
type="uri_file",确保Azure ML正确识别输入类型并解析资产引用。
内容的提问来源于stack exchange,提问作者WEVERETT
相关产品推荐
相关产品推荐

