You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML v2 Pipeline:无需虚设输入输出设置作业执行顺序

Azure ML Pipeline 按顺序执行脚本的简化配置方法

我有三个Python脚本a.py、b.py、c.py,需要在Azure ML Studio Pipeline中按顺序执行。尝试编写了如下代码,但无法实现作业的顺序运行。需要在Azure ML Pipelines中分步执行而非将三个脚本作为单个作业运行,已知可通过创建虚设输入输出来设置作业层级,但觉得该方式较为复杂,想询问是否存在更简便的方式(例如类似b_job.run_after(a_job)的配置方法)来设置作业执行顺序。

原代码

import warnings
warnings.filterwarnings("ignore")
import yaml
from azure.ai.ml import command, dsl
from azure.ai.ml.entities import PipelineJobSettings
import os
import sys
# Append the directory to system path
sys.path.append(os.path.join(os.path.dirname(__file__), ".."))

# Load environment variables
if __name__ == "__main__":
    from dotenv import load_dotenv, find_dotenv

    load_dotenv(find_dotenv())

    # Load configuration from YAML file
    with open(os.path.join("conda.yaml"), encoding="utf-8") as stream:
        config = yaml.safe_load(stream)

    # Initialize MLStudioHandler
    from ml_studio_jobs.mlstudio_handling import MLStudioHandler

    ml_studio_handler = MLStudioHandler()
    env = ml_studio_handler.get_env_version(config['name'])
    compute = ml_studio_handler.get_or_start_compute(os.environ.get("COMPUTE_NAME"))

    mode = "a"
    a_job = command(
        code="/.",  # location of source code
        command="python a.py",
        environment=env,
        display_name=f"test_{mode}",
    )
    a_component = ml_studio_handler.create_or_update(a_job.component)

    mode = "b"
    b_job = command(
        code="/.",  # location of source code
        command="python b.py",
        environment=env,
        display_name=f"test_{mode}",
    )
    b_component = ml_studio_handler.create_or_update(b_job.component)

    mode = "c"
    c_job = command(
        code="/.",  # location of source code
        command="python c.py",
        environment=env,
        display_name=f"test_{mode}",
    )
    c_component = ml_studio_handler.create_or_update(c_job.component)


    # Define the pipeline
    @dsl.pipeline(
        compute=compute.name,
        description="E2E Churn Prediction Pipeline",
    )
    def process():
        # Step 1: Get Data - produces dummy output
        a_job = a_component()
        b_job = b_component()
        c_job = c_component()


    # Create the pipeline
    pipeline = process()
    pipeline.settings = PipelineJobSettings(force_rerun=True, continue_on_step_failure=True)

    # Submit the pipeline job
    pipeline_job = ml_studio_handler.create_or_update_jobs(
        jobs=pipeline,
        experiment_name="test_experiment_ml",
    )

简化的顺序配置方法

原代码中三个作业实例未设置依赖关系,因此会并行执行。Azure ML v2 SDK提供了两种无需虚设输入输出的简化配置方式:

方式1:使用depends_on参数

在创建组件实例时,通过depends_on参数指定依赖的作业列表:

@dsl.pipeline(
    compute=compute.name,
    description="E2E Churn Prediction Pipeline",
)
def process():
    a_job = a_component()
    # 指定b_job依赖a_job完成后执行
    b_job = b_component(depends_on=[a_job])
    # 指定c_job依赖b_job完成后执行
    c_job = c_component(depends_on=[b_job])

方式2:使用.after()方法

通过调用作业实例的.after()方法,直接指定依赖的前置作业:

@dsl.pipeline(
    compute=compute.name,
    description="E2E Churn Prediction Pipeline",
)
def process():
    a_job = a_component()
    b_job = b_component()
    # 设置b_job在a_job之后执行
    b_job.after(a_job)
    c_job = c_component()
    # 设置c_job在b_job之后执行
    c_job.after(b_job)

两种方式都能实现a.py → b.py → c.py的顺序执行,无需额外创建虚设输入输出,配置更直观简洁。

内容的提问来源于stack exchange,提问作者Marios

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 14:02:06