如何在Azure ML中基于触发时间设置数据集调度并适配格式?
解决Azure ML调度触发时间格式适配数据集路径的方案
方法1:直接用调度表达式格式化参数
Azure ML调度支持对${{creation_context.trigger_time}}使用内置的date()函数做格式化,直接生成符合路径要求的字符串。
CLI配置示例
在调度的YAML配置里直接拼接路径参数:
schedules: - name: daily_data_schedule trigger: type: cron cron: "0 0 * * *" create_job: type: pipeline pipeline: azureml:your-pipeline:1 parameters: data_path: "path_on_datastore/${{creation_context.trigger_time | date('yyyy/MM/dd')}}/some_data.tsv"
Python SDK示例
创建调度时直接传入格式化后的参数:
from azure.ai.ml import MLClient, Schedule, ScheduleRecurrence from azure.identity import DefaultAzureCredential ml_client = MLClient(DefaultAzureCredential(), "订阅ID", "资源组名", "工作区名") schedule = Schedule( name="daily_data_schedule", trigger=ScheduleRecurrence(frequency="Day", interval=1), create_job=dict( type="pipeline", pipeline="azureml:your-pipeline:1", parameters={ "data_path": "path_on_datastore/${{creation_context.trigger_time | date('yyyy/MM/dd')}}/some_data.tsv" } ) ) ml_client.schedules.begin_create_or_update(schedule)
这里的date('yyyy/MM/dd')会把触发时间转换成你需要的年/月/日路径格式。
方法2:用轻量Python组件处理复杂格式
如果需要更灵活的时间处理(比如日期偏移、自定义格式拼接),可以在Pipeline里加一个简单的Python组件,输入触发时间参数,输出格式化后的路径字符串,再传递给后续组件。
1. 定义时间格式化组件
# format_time.py def format_trigger_time(trigger_time: str, base_path: str = "path_on_datastore") -> str: from datetime import datetime # 解析Azure ML返回的ISO格式触发时间(如'2023-01-01T00:00:00Z') dt = datetime.fromisoformat(trigger_time.replace('Z', '+00:00')) # 生成目标路径格式 date_segment = dt.strftime("%Y/%m/%d") return f"{base_path}/{date_segment}/some_data.tsv"
2. 在Pipeline中调用组件
from azure.ai.ml import load_component, Pipeline # 加载自定义组件和数据处理组件 format_component = load_component(source="format_time_component.yaml") data_process_component = load_component(source="data_process_component.yaml") pipeline = Pipeline() # 第一步:格式化时间生成路径 format_step = pipeline.add_job( format_component, inputs=dict(trigger_time="${{creation_context.trigger_time}}") ) # 第二步:用格式化后的路径读取数据 data_step = pipeline.add_job( data_process_component, inputs=dict(data_path=format_step.outputs.formatted_path) )
方法3:给数据集配置动态路径参数
如果使用Azure ML的FileDataset或TabularDataset,可以直接在数据集定义中设置动态参数,结合调度触发时间自动生成路径:
from azure.ai.ml.entities import FileDataset # 关联数据存储并设置动态路径参数 dataset = FileDataset( path=[(datastore, "path_on_datastore/{year}/{month}/{day}/some_data.tsv")], parameters={ "year": "${{creation_context.trigger_time | date('yyyy')}}", "month": "${{creation_context.trigger_time | date('MM')}}", "day": "${{creation_context.trigger_time | date('dd')}}" } )
调度触发时,会自动替换参数生成对应日期的数据集路径。
内容的提问来源于stack exchange,提问作者leuction
相关产品推荐
相关产品推荐

