如何在Azure ML批量端点管道中添加并行运行步骤?
问题解答:Azure ML批量端点管道中添加并行步骤的配置修正
可以在Azure ML批量端点管道中加入并行运行步骤,部署失败源于配置中的输入输出映射不匹配、参数引用错误等问题,以下是具体修改方案:
一、修正pipeline.yaml的输入输出映射错误
- 对齐pipeline输入与并行步骤的引用名称
- 统一并行步骤输出与后续命令步骤的引用名称
- 修正pipeline输出定义与后续步骤的名称一致性
修改后的pipeline.yaml:
$schema: https://azuremlschemas.azureedge.net/latest/pipelinecomponent.schema.json type: pipeline name: pipeline inputs: data: type: string outputs: scores: type: uri_folder mode: upload jobs: process_data: type: parallel component: process_data.yaml inputs: data: ${{parent.inputs.data}} model_file: azureml:<model_name> trigger_next_step: type: command component: trigger_next_step.yaml inputs: data: ${{parent.jobs.process_data.outputs.output_folder}} outputs: scores: mode: upload path: ${{parent.outputs.scores}}
二、完善parallel组件配置(process_data.yaml)
确保组件参数定义完整,避免因隐式缺失导致的键值对错误:
$schema: https://azuremlschemas.azureedge.net/latest/parallelComponent.schema.json type: parallel name: parallel_embed description: parallel embed display_name: parallel_embedding compute: azureml:reip-etl-mip input_data: ${{inputs.data}} inputs: model_file: type: mlflow_model description: Trained MLflow model outputs: output_folder: mode: rw_mount type: uri_folder description: Output folder for processed data resources: instance_count: 1 max_concurrency_per_instance: 1 logging_level: "INFO" mini_batch_size: '1' task: type: run_function code: "./" entry_script: batch_score.py environment: conda_file: environment.yaml image: mcr.microsoft.com/azureml/openmpi3.1.2-ubuntu18.04 program_arguments: >- --model_file ${{inputs.model_file}} --output_folder ${{outputs.output_folder}}
三、批量部署配置的验证(batch_deployment.yaml)
确认默认计算集群名称正确,无需额外修改,仅需替换占位符:
$schema: https://azuremlschemas.azureedge.net/latest/pipelineComponentBatchDeployment.schema.json name: pipeline-test endpoint_name: ptest type: pipeline component: pipeline.yaml settings: continue_on_step_failure: true default_compute: <your-actual-compute-cluster>
报错原因说明
你遇到的Value cannot be null. (Parameter 'key')错误,核心是输入输出的名称引用不匹配:比如pipeline定义的输入为data,但并行步骤错误引用了不存在的input_val;并行组件输出为output_folder,后续步骤却引用了未定义的output_val,导致系统无法识别有效参数键值对,从而抛出空值错误。
内容的提问来源于stack exchange,提问作者WEVERETT
相关产品推荐
相关产品推荐

