You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML管道作业中MLTable文件路径配置及报错解决

解决Azure ML中MLTable路径无效的报错

问题背景

参考azureml-examples仓库的《automl-forecasting-task-energy-demand-advanced》笔记本,本地创建MLTable文件时使用相对路径引用CSV,配置的Python输入代码如下:

my_training_data_input = Input(
    type=AssetTypes.MLTABLE, path="./data/training-mltable-folder"
)

compute = AmlCompute(
        name=compute_name, size="STANDARD_D2_V2", min_instances=0, max_instances=4
    )

forecasting_job = automl.forecasting(
    compute=compute_name, # name of the compute target we created above
    # name="dpv2-forecasting-job-02",
    experiment_name=exp_name,
    training_data=my_training_data_input,
    # validation_data = my_validation_data_input,
    target_column_name="demand",
    primary_metric="NormalizedRootMeanSquaredError",
    n_cross_validations="auto",
    enable_model_explainability=True,
    tags={"my_custom_tag": "My custom value"},
)

returned_job = ml_client.jobs.create_or_update(
    forecasting_job
)

ml_client.jobs.stream(returned_job.name)

(注:原代码中path=""./data/training-mltable-folder"存在语法错误,已修正为path="./data/training-mltable-folder")

报错信息

Encountered user error while fetching data from Dataset. Error: UserErrorException:
Message: MLTable yaml schema is invalid:
Error Code: Validation
Validation Error Code: Invalid MLTable
Validation Target: MLTableToDataflow
Error Message: Failed to convert a MLTable to dataflow
uri path is not a valid datastore uri path
| session_id=857bd9a1-097b-4df6-aa1c-8871f89580d8
InnerException None
ErrorResponse
{
"error": {
"code": "UserError",
"message": "MLTable yaml schema is invalid: \nError Code: Validation\nValidation Error Code: Invalid MLTable\nValidation Target: MLTableToDataflow\nError Message: Failed to convert a MLTable to dataflow\nuri path is not a valid datastore uri path\n| session_id=857bd9a1-097b-4df6-aa1c-8871f89580d8"
}
}

当前MLTable配置

paths:
  - file: ./nyc_energy_training_clean.csv
transformations:
  - read_delimited:
        delimiter: ','
        encoding: 'ascii'
  - convert_column_types:
      - columns: demand
        column_type: float
      - columns: precip
        column_type: float
      - columns: temp
        column_type: float

解决方案

报错核心原因:提交到Azure ML远程计算集群的作业无法直接访问本地文件路径,必须将数据上传到Azure ML数据存储,并使用符合平台要求的路径格式引用。

步骤1:上传本地数据到Azure ML数据存储

将包含MLTable文件和CSV的./data/training-mltable-folder文件夹上传并注册为MLTable数据资产,代码示例:

from azure.ai.ml.entities import Data
from azure.ai.ml.constants import AssetTypes

# 定义MLTable数据资产
my_mltable = Data(
    path="./data/training-mltable-folder",
    type=AssetTypes.MLTABLE,
    name="nyc-energy-training-mltable",
    description="NYC energy demand training data as MLTable",
)

# 上传并注册到工作区
ml_client.data.create_or_update(my_mltable)

步骤2:调整MLTable路径配置

上传完成后,MLTable中的路径可使用相对路径(相对于MLTable文件所在的文件夹),保持原有配置即可:

paths:
  - file: nyc_energy_training_clean.csv
# 其余转换配置不变

也可使用数据存储绝对路径格式(可选):

paths:
  - file: azureml://datastores/workspaceblobstore/paths/data/training-mltable-folder/nyc_energy_training_clean.csv

步骤3:修改Input引用方式

在Python代码中,使用注册后的MLTable资产名称而非本地路径引用数据:

my_training_data_input = Input(
    type=AssetTypes.MLTABLE, path="azureml:nyc-energy-training-mltable:1"
)

其中1为资产版本号,需根据实际注册的版本调整。

额外说明

  • 本地调试作业时,需确保MLTable与CSV的相对路径正确且运行环境可访问本地文件;但提交到远程计算集群时,必须依赖Azure ML数据存储中的数据。
  • 禁止在MLTable中使用本地绝对路径,该路径在远程计算环境中完全无效。

内容的提问来源于stack exchange,提问作者MrFranzén

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 22:32:22