You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Data Factory:如何强制刷新Azure Machine Learning管道?

强制刷新Azure ML管道获取最新结果的方法
  • 禁用管道运行缓存
    直接在管道步骤定义里设置allow_reuse=False,让该步骤每次都重新执行,不会复用之前的缓存结果:

    from azureml.pipeline.core import PipelineStep
    
    step = PipelineStep(
        name="your_step_name",
        script_name="script.py",
        compute_target=compute_target,
        allow_reuse=False
    )
    

    要是想让整个管道都不复用缓存,提交运行时加reuse_cache=False参数:

    pipeline_run = experiment.submit(pipeline, reuse_cache=False)
    
  • 修改输入输出触发重新运行
    AML管道的缓存是基于输入输出的哈希值判断的,只要改动输入数据(比如给数据源加时间戳标记更新、替换数据集文件)或者调整输出路径,就能打破缓存,触发全新运行。比如给数据集添加时间戳列:

    from azureml.core import Dataset
    
    input_data = Dataset.get_by_name(workspace, name="your_dataset").with_timestamp_column("last_modified")
    
  • 用动态参数强制刷新
    添加一个动态生成的参数(比如当前时间戳),每次提交管道时传入不同的值,让系统认为输入有变化,从而跳过缓存:

    from azureml.pipeline.core.graph import PipelineParameter
    from datetime import datetime
    
    timestamp_param = PipelineParameter(name="timestamp", default_value=str(datetime.now()))
    step = PipelineStep(
        name="your_step_name",
        script_name="script.py",
        arguments=["--timestamp", timestamp_param],
        compute_target=compute_target
    )
    
  • 在Azure Data Factory中配置触发参数
    如果是通过ADF触发AML管道,在ADF的AML管道活动设置里,添加reuse_cache=False的运行参数,或者传入动态生成的时间戳,确保每次触发的参数都不一样,强制AML重新执行管道。

内容的提问来源于stack exchange,提问作者DataWrangler1980

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 05:42:08