如何在CI/CD流程中更新已调度的AzureML管道?
AzureML调度管道的CI/CD更新优化方案
不需要每次都禁用并重建管道和调度,有两种更高效的方式实现更新:
方式一:发布新版本管道并更新调度关联
每次更新管道逻辑后,以同名发布新的管道版本,再将现有调度关联到新版本即可,这种方式会保留历史版本,方便后续回滚。
示例代码(中文注释)
from azureml.core import Workspace, Pipeline, Schedule from azureml.pipeline.core.schedule import ScheduleRecurrence # 初始化AzureML工作区 ws = Workspace( subscription_id="你的订阅ID", resource_group="你的资源组", workspace_name="你的工作区名称", ) # --- 首次创建管道与调度 --- # 定义初始管道 pipeline = Pipeline(workspace=ws, steps=...) # 发布管道,指定固定名称 published_pipeline = pipeline.publish( name="业务处理管道", description="初始版本:基础数据处理逻辑" ) # 创建每日执行的调度 my_recurrence = ScheduleRecurrence(frequency="Day", interval=1) schedule = Schedule.create( workspace=ws, name="每日数据处理调度", pipeline_id=published_pipeline.id, experiment_name="数据处理实验", recurrence=my_recurrence, description="每日凌晨执行数据处理管道" ) # --- 后续更新管道逻辑 --- # 定义更新后的新管道 new_pipeline = Pipeline(workspace=ws, steps=...) # 以同名发布新版本,系统会自动生成新的版本号 new_published_pipeline = new_pipeline.publish( name="业务处理管道", # 必须与原有已发布管道同名 description="更新版本:新增异常数据过滤逻辑" ) # 找到现有调度并更新关联的管道ID existing_schedule = Schedule.get(ws, name="每日数据处理调度") existing_schedule.update(pipeline_id=new_published_pipeline.id)
方式二:直接更新已发布管道的定义
如果不需要保留历史版本,可以直接覆盖原有已发布管道的定义,操作更简洁,但无法回滚到旧版本,适合测试环境快速迭代。
示例代码(中文注释)
from azureml.core import Workspace, Pipeline from azureml.pipeline.core import PublishedPipeline ws = Workspace( subscription_id="你的订阅ID", resource_group="你的资源组", workspace_name="你的工作区名称", ) # 定义更新后的新管道 new_pipeline = Pipeline(workspace=ws, steps=...) # 序列化新管道的定义 new_pipeline_def = new_pipeline._serialize() # 获取已发布的旧管道对象 existing_published_pipeline = PublishedPipeline.get(ws, id="原有已发布管道ID") # 直接更新管道定义 existing_published_pipeline.update(definition=new_pipeline_def)
两种方式对比
- 版本化更新:保留所有历史版本,支持回滚,符合生产环境的合规要求,推荐优先使用。
- 直接更新:无版本历史,操作步骤少,适合测试阶段快速验证逻辑。
内容的提问来源于stack exchange,提问作者techtech
相关产品推荐
相关产品推荐

