Azure ML管道作业中MLTable文件路径配置及报错解决
解决Azure ML中MLTable路径无效的报错
问题背景
参考azureml-examples仓库的《automl-forecasting-task-energy-demand-advanced》笔记本,本地创建MLTable文件时使用相对路径引用CSV,配置的Python输入代码如下:
my_training_data_input = Input( type=AssetTypes.MLTABLE, path="./data/training-mltable-folder" ) compute = AmlCompute( name=compute_name, size="STANDARD_D2_V2", min_instances=0, max_instances=4 ) forecasting_job = automl.forecasting( compute=compute_name, # name of the compute target we created above # name="dpv2-forecasting-job-02", experiment_name=exp_name, training_data=my_training_data_input, # validation_data = my_validation_data_input, target_column_name="demand", primary_metric="NormalizedRootMeanSquaredError", n_cross_validations="auto", enable_model_explainability=True, tags={"my_custom_tag": "My custom value"}, ) returned_job = ml_client.jobs.create_or_update( forecasting_job ) ml_client.jobs.stream(returned_job.name)
(注:原代码中path=""./data/training-mltable-folder"存在语法错误,已修正为path="./data/training-mltable-folder")
报错信息
Encountered user error while fetching data from Dataset. Error: UserErrorException: Message: MLTable yaml schema is invalid: Error Code: Validation Validation Error Code: Invalid MLTable Validation Target: MLTableToDataflow Error Message: Failed to convert a MLTable to dataflow uri path is not a valid datastore uri path | session_id=857bd9a1-097b-4df6-aa1c-8871f89580d8 InnerException None ErrorResponse { "error": { "code": "UserError", "message": "MLTable yaml schema is invalid: \nError Code: Validation\nValidation Error Code: Invalid MLTable\nValidation Target: MLTableToDataflow\nError Message: Failed to convert a MLTable to dataflow\nuri path is not a valid datastore uri path\n| session_id=857bd9a1-097b-4df6-aa1c-8871f89580d8" } }
当前MLTable配置
paths: - file: ./nyc_energy_training_clean.csv transformations: - read_delimited: delimiter: ',' encoding: 'ascii' - convert_column_types: - columns: demand column_type: float - columns: precip column_type: float - columns: temp column_type: float
解决方案
报错核心原因:提交到Azure ML远程计算集群的作业无法直接访问本地文件路径,必须将数据上传到Azure ML数据存储,并使用符合平台要求的路径格式引用。
步骤1:上传本地数据到Azure ML数据存储
将包含MLTable文件和CSV的./data/training-mltable-folder文件夹上传并注册为MLTable数据资产,代码示例:
from azure.ai.ml.entities import Data from azure.ai.ml.constants import AssetTypes # 定义MLTable数据资产 my_mltable = Data( path="./data/training-mltable-folder", type=AssetTypes.MLTABLE, name="nyc-energy-training-mltable", description="NYC energy demand training data as MLTable", ) # 上传并注册到工作区 ml_client.data.create_or_update(my_mltable)
步骤2:调整MLTable路径配置
上传完成后,MLTable中的路径可使用相对路径(相对于MLTable文件所在的文件夹),保持原有配置即可:
paths: - file: nyc_energy_training_clean.csv # 其余转换配置不变
也可使用数据存储绝对路径格式(可选):
paths: - file: azureml://datastores/workspaceblobstore/paths/data/training-mltable-folder/nyc_energy_training_clean.csv
步骤3:修改Input引用方式
在Python代码中,使用注册后的MLTable资产名称而非本地路径引用数据:
my_training_data_input = Input( type=AssetTypes.MLTABLE, path="azureml:nyc-energy-training-mltable:1" )
其中1为资产版本号,需根据实际注册的版本调整。
额外说明
- 本地调试作业时,需确保MLTable与CSV的相对路径正确且运行环境可访问本地文件;但提交到远程计算集群时,必须依赖Azure ML数据存储中的数据。
- 禁止在MLTable中使用本地绝对路径,该路径在远程计算环境中完全无效。
内容的提问来源于stack exchange,提问作者MrFranzén
相关产品推荐
相关产品推荐

