You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML管道构建问题:OutputFileDatasetConfig与Datastore路径配置

问题分析与解决方案

你遇到的核心问题是混淆了OutputFileDatasetConfig的管道输出机制和Dataset.Tabular.register_pandas_dataframe的用法,同时存在参数笔误。以下是具体修正方案:

1. 修正dataprep.py代码

# -----dataprep stuff and imports
import argparse
import pandas as pd
from azureml.core import Run, Dataset
from azureml.data.datapath import DataPath

parser = argparse.ArgumentParser()
parser.add_argument("--month_train", required=True)
parser.add_argument("--year_train", required=True)
parser.add_argument('--output_path', dest='output_path', required=True)

args = parser.parse_args()

run = Run.get_context()
ws = run.experiment.workspace
datastore = ws.get_default_datastore()

name_dataset_input = 'Customer_data_' + str(args.year_train)
name_dataset_output = 'DATA_PREP_' + str(args.year_train) + '_' + str(args.month_train)

# 获取输入数据集
ds = Dataset.get_by_name(ws, name_dataset_input)
df = ds.to_pandas_dataframe()

# 修正笔误:mois_train -> month_train
df = apply(df, args.month_train)

# 第一步:将处理后的数据写入管道输出指定的本地路径
# OutputFileDatasetConfig会自动将本地文件同步到Datastore
output_csv_path = f"{args.output_path}/prepped_data.csv"
df.to_csv(output_csv_path, index=False)

# 第二步:从Datastore路径创建并注册TabularDataset
datastore_file_path = DataPath(datastore, f"managed-dataset/{run.id}/output_path/prepped_data.csv")
prepped_dataset = Dataset.Tabular.from_delimited_files(path=datastore_file_path)
prepped_dataset.register(workspace=ws, name=name_dataset_output, create_new_version=True)

2. 修正管道定义代码

from azureml.data import OutputFileDatasetConfig
from azureml.pipeline.steps import PythonScriptStep

prepped_data_path = OutputFileDatasetConfig(
    name="output_path",
    destination=(datastore, 'managed-dataset/{run-id}/{output-name}'),
    # 指定上传文件名,确保Datastore中文件路径可预测
    upload_file_name="prepped_data.csv"
)

dataprep_step = PythonScriptStep(
    name="dataprep", 
    script_name="dataprep.py", 
    compute_target=compute_target, 
    runconfig=aml_run_config,
    arguments=[
        "--output_path", prepped_data_path, 
        "--month_train", month_train,
        "--year_train", year_train
    ],
    # 必须声明输出,否则后续步骤无法引用
    outputs=[prepped_data_path],
    allow_reuse=True
)

关键修正点说明

  • OutputFileDatasetConfig的正确用法:它会在计算目标上映射一个本地路径,写入该路径的文件会自动同步到你指定的Datastore目录。你需要先将DataFrame写入这个本地路径,而非直接用register_pandas_dataframe。
  • 注册数据集的正确方式:register_pandas_dataframe需要传入Datastore对象和相对路径,而非管道传递的本地路径引用。我们先写入本地路径,再从同步后的Datastore路径创建数据集并注册。
  • 管道输出声明:必须在PythonScriptStep中添加outputs=[prepped_data_path],后续AutoML步骤可通过prepped_data_path.as_input(name="training_data")直接引用该输出。
  • 笔误修正:代码中args.mois_train是拼写错误,应改为args.month_train,否则会导致参数未定义错误。

后续AutoML步骤引用示例

from azureml.pipeline.steps import AutoMLStep

automl_step = AutoMLStep(
    name="automl_train",
    automl_config=automl_config,
    inputs=[prepped_data_path.as_input(name="training_data")],
    compute_target=compute_target,
    allow_reuse=True
)

内容的提问来源于stack exchange,提问作者Mavil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 09:57:30