You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用YAML版RunConfiguration运行AzureML管道时遇pandas模块找不到错误

Azure ML Pipeline执行时出现Pandas ModuleNotFoundError

问题

编排Azure Machine Learning Pipeline时,执行阶段抛出ModuleNotFoundError: No module named 'pandas'。单独运行脚本T01_Test_Task.py正常,但放入管道后无法执行。

相关代码与配置

管道编排代码

# Loading run config
print("Loading run config")
task_1_run_config = RunConfiguration.load(
    os.path.join(WORKING_DIR + '/pipeline/task_runconfigs/T01_Test_Task.yml')
    ) 

task_1_script_run_config = ScriptRunConfig(
    source_directory=os.path.join(WORKING_DIR + '/pipeline/task_scripts'),
    run_config=task_1_run_config    
)

task_1_py_script_step = PythonScriptStep(
    name='Task_1_Step',
    script_name=task_1_script_run_config.script,
    source_directory=task_1_script_run_config.source_directory,
    compute_target=compute_target
)

pipeline_run_config = Pipeline(workspace=workspace, steps=[task_1_py_script_step])#, task_2])

pipeline_run = Experiment(workspace, 'Test_Run_New_Pipeline').submit(pipeline_run_config)
pipeline_run.wait_for_completion()

environment.yml配置

name: phinmo_pipeline_env
dependencies:
- python=3.8
- pip:
  - pandas
  - azureml-core==1.43.0
  - azureml-sdk
  - scipy
  - scikit-learn
  - numpy
  - pyyaml==6.0
  - datetime
  - azure
channels:
  - conda-forge

RunConfiguration文件(T01_Test_Task.yml)

# The script to run.
script: T01_Test_Task.py
# The arguments to the script file.
arguments: [
  "--test", False,
  "--date", "2022-07-26"
]
# The name of the compute target to use for this run.
compute_target: phinmo-compute-cluster
# Framework to execute inside. Allowed values are "Python", "PySpark", "CNTK", "TensorFlow", and "PyTorch".
framework: Python
# Maximum allowed duration for the run.
maxRunDurationSeconds: 6000
# Number of nodes to use for running job.
nodeCount: 1

#Environment details.
environment:
  # Environment name
  name: phinmo_pipeline_env
  # Environment version
  version:
  # Environment variables set for the run.
  #environmentVariables:
  #  EXAMPLE_ENV_VAR: EXAMPLE_VALUE
  # Python details
  python:
    # user_managed_dependencies=True indicates that the environmentwill be user managed. False indicates that AzureML willmanage the user environment.
    userManagedDependencies: false
    # The python interpreter path
    interpreterPath: python
    # Path to the conda dependencies file to use for this run. If a project
    # contains multiple programs with different sets of dependencies, it may be
    # convenient to manage those environments with separate files.
    condaDependenciesFile: environment.yml
    # The base conda environment used for incremental environment creation.
    baseCondaEnvironment: AzureML-sklearn-0.24-ubuntu18.04-py37-cpu
  # Docker details
  
# History details.
history:
  # Enable history tracking -- this allows status, logs, metrics, and outputs
  # to be collected for a run.
  outputCollection: true
  # Whether to take snapshots for history.
  snapshotProject: true
  # Directories to sync with FileWatcher.
  directoriesToWatch:
  - logs
# data reference configuration details
dataReferences: {}
# The configuration details for data.
data: {}
# Project share datastore reference.
sourceDirectoryDataStore:

已尝试方案

  • 使用environment.python.conda_dependencies对象覆盖RunConfiguration的environment属性
  • 在environment.yml中指定pandas具体版本
  • 调整environment.yml的存放位置

可能的解决方法

1. 修复Python版本冲突

environment.yml指定python=3.8,但RunConfig中baseCondaEnvironment使用的是AzureML-sklearn-0.24-ubuntu18.04-py37-cpu(基于Python3.7),版本冲突会导致conda无法正确安装依赖。

  • 方案:将base环境替换为Python3.8版本,比如AzureML-sklearn-0.24-ubuntu18.04-py38-cpu;或把environment.yml中的python=3.8改为python=3.7。

2. 确认condaDependenciesFile路径正确

RunConfig中condaDependenciesFile: environment.yml,但ScriptRunConfig的source_directory是/pipeline/task_scripts,如果environment.yml实际在/pipeline/task_runconfigs目录,需要调整路径为相对路径:

condaDependenciesFile: ../task_runconfigs/environment.yml

或者将environment.yml移动到task_scripts目录下。

3. 手动注册环境到Workspace

提前创建并注册环境,避免管道运行时环境创建出错:

from azureml.core import Environment

# 从conda文件创建环境
env = Environment.from_conda_specification(
    name="phinmo_pipeline_env", 
    file_path=os.path.join(WORKING_DIR, '/pipeline/task_runconfigs/environment.yml')
)
# 注册到工作区
env.register(workspace=workspace)

# 替换RunConfig中的环境
task_1_run_config.environment = env

4. 禁用环境缓存,强制重建

AzureML可能缓存旧环境,导致新依赖未安装。在PythonScriptStep中添加allow_reuse=False:

task_1_py_script_step = PythonScriptStep(
    name='Task_1_Step',
    script_name=task_1_script_run_config.script,
    source_directory=task_1_script_run_config.source_directory,
    compute_target=compute_target,
    allow_reuse=False
)

或提交管道时强制重新生成输出:

pipeline_run = Experiment(workspace, 'Test_Run_New_Pipeline').submit(
    pipeline_run_config, 
    regenerate_outputs=True
)

5. 查看环境构建日志排查安装错误

在AzureML Studio中找到对应的运行任务,查看azureml-logs/20_image_build_log.txt,里面会记录环境创建的详细过程,确认pandas是否被正确安装,是否有安装失败的报错信息。


内容的提问来源于stack exchange,提问作者nils.hahn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 05:24:10