You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Azure ML中基于触发时间设置数据集调度并适配格式?

解决Azure ML调度触发时间格式适配数据集路径的方案

方法1:直接用调度表达式格式化参数

Azure ML调度支持对${{creation_context.trigger_time}}使用内置的date()函数做格式化,直接生成符合路径要求的字符串。

CLI配置示例

在调度的YAML配置里直接拼接路径参数:

schedules:
  - name: daily_data_schedule
    trigger:
      type: cron
      cron: "0 0 * * *"
    create_job:
      type: pipeline
      pipeline: azureml:your-pipeline:1
      parameters:
        data_path: "path_on_datastore/${{creation_context.trigger_time | date('yyyy/MM/dd')}}/some_data.tsv"

Python SDK示例

创建调度时直接传入格式化后的参数:

from azure.ai.ml import MLClient, Schedule, ScheduleRecurrence
from azure.identity import DefaultAzureCredential

ml_client = MLClient(DefaultAzureCredential(), "订阅ID", "资源组名", "工作区名")

schedule = Schedule(
    name="daily_data_schedule",
    trigger=ScheduleRecurrence(frequency="Day", interval=1),
    create_job=dict(
        type="pipeline",
        pipeline="azureml:your-pipeline:1",
        parameters={
            "data_path": "path_on_datastore/${{creation_context.trigger_time | date('yyyy/MM/dd')}}/some_data.tsv"
        }
    )
)

ml_client.schedules.begin_create_or_update(schedule)

这里的date('yyyy/MM/dd')会把触发时间转换成你需要的年/月/日路径格式。

方法2:用轻量Python组件处理复杂格式

如果需要更灵活的时间处理(比如日期偏移、自定义格式拼接),可以在Pipeline里加一个简单的Python组件,输入触发时间参数,输出格式化后的路径字符串,再传递给后续组件。

1. 定义时间格式化组件

# format_time.py
def format_trigger_time(trigger_time: str, base_path: str = "path_on_datastore") -> str:
    from datetime import datetime
    # 解析Azure ML返回的ISO格式触发时间(如'2023-01-01T00:00:00Z')
    dt = datetime.fromisoformat(trigger_time.replace('Z', '+00:00'))
    # 生成目标路径格式
    date_segment = dt.strftime("%Y/%m/%d")
    return f"{base_path}/{date_segment}/some_data.tsv"

2. 在Pipeline中调用组件

from azure.ai.ml import load_component, Pipeline

# 加载自定义组件和数据处理组件
format_component = load_component(source="format_time_component.yaml")
data_process_component = load_component(source="data_process_component.yaml")

pipeline = Pipeline()
# 第一步:格式化时间生成路径
format_step = pipeline.add_job(
    format_component,
    inputs=dict(trigger_time="${{creation_context.trigger_time}}")
)
# 第二步:用格式化后的路径读取数据
data_step = pipeline.add_job(
    data_process_component,
    inputs=dict(data_path=format_step.outputs.formatted_path)
)

方法3:给数据集配置动态路径参数

如果使用Azure ML的FileDataset或TabularDataset,可以直接在数据集定义中设置动态参数,结合调度触发时间自动生成路径:

from azure.ai.ml.entities import FileDataset

# 关联数据存储并设置动态路径参数
dataset = FileDataset(
    path=[(datastore, "path_on_datastore/{year}/{month}/{day}/some_data.tsv")],
    parameters={
        "year": "${{creation_context.trigger_time | date('yyyy')}}",
        "month": "${{creation_context.trigger_time | date('MM')}}",
        "day": "${{creation_context.trigger_time | date('dd')}}"
    }
)

调度触发时,会自动替换参数生成对应日期的数据集路径。


内容的提问来源于stack exchange,提问作者leuction

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 17:08:11