使用YAML版RunConfiguration运行AzureML管道时遇pandas模块找不到错误
问题
编排Azure Machine Learning Pipeline时,执行阶段抛出ModuleNotFoundError: No module named 'pandas'。单独运行脚本T01_Test_Task.py正常,但放入管道后无法执行。
相关代码与配置
管道编排代码
# Loading run config print("Loading run config") task_1_run_config = RunConfiguration.load( os.path.join(WORKING_DIR + '/pipeline/task_runconfigs/T01_Test_Task.yml') ) task_1_script_run_config = ScriptRunConfig( source_directory=os.path.join(WORKING_DIR + '/pipeline/task_scripts'), run_config=task_1_run_config ) task_1_py_script_step = PythonScriptStep( name='Task_1_Step', script_name=task_1_script_run_config.script, source_directory=task_1_script_run_config.source_directory, compute_target=compute_target ) pipeline_run_config = Pipeline(workspace=workspace, steps=[task_1_py_script_step])#, task_2]) pipeline_run = Experiment(workspace, 'Test_Run_New_Pipeline').submit(pipeline_run_config) pipeline_run.wait_for_completion()
environment.yml配置
name: phinmo_pipeline_env dependencies: - python=3.8 - pip: - pandas - azureml-core==1.43.0 - azureml-sdk - scipy - scikit-learn - numpy - pyyaml==6.0 - datetime - azure channels: - conda-forge
RunConfiguration文件(T01_Test_Task.yml)
# The script to run. script: T01_Test_Task.py # The arguments to the script file. arguments: [ "--test", False, "--date", "2022-07-26" ] # The name of the compute target to use for this run. compute_target: phinmo-compute-cluster # Framework to execute inside. Allowed values are "Python", "PySpark", "CNTK", "TensorFlow", and "PyTorch". framework: Python # Maximum allowed duration for the run. maxRunDurationSeconds: 6000 # Number of nodes to use for running job. nodeCount: 1 #Environment details. environment: # Environment name name: phinmo_pipeline_env # Environment version version: # Environment variables set for the run. #environmentVariables: # EXAMPLE_ENV_VAR: EXAMPLE_VALUE # Python details python: # user_managed_dependencies=True indicates that the environmentwill be user managed. False indicates that AzureML willmanage the user environment. userManagedDependencies: false # The python interpreter path interpreterPath: python # Path to the conda dependencies file to use for this run. If a project # contains multiple programs with different sets of dependencies, it may be # convenient to manage those environments with separate files. condaDependenciesFile: environment.yml # The base conda environment used for incremental environment creation. baseCondaEnvironment: AzureML-sklearn-0.24-ubuntu18.04-py37-cpu # Docker details # History details. history: # Enable history tracking -- this allows status, logs, metrics, and outputs # to be collected for a run. outputCollection: true # Whether to take snapshots for history. snapshotProject: true # Directories to sync with FileWatcher. directoriesToWatch: - logs # data reference configuration details dataReferences: {} # The configuration details for data. data: {} # Project share datastore reference. sourceDirectoryDataStore:
已尝试方案
- 使用
environment.python.conda_dependencies对象覆盖RunConfiguration的environment属性 - 在environment.yml中指定pandas具体版本
- 调整environment.yml的存放位置
可能的解决方法
1. 修复Python版本冲突
environment.yml指定python=3.8,但RunConfig中baseCondaEnvironment使用的是AzureML-sklearn-0.24-ubuntu18.04-py37-cpu(基于Python3.7),版本冲突会导致conda无法正确安装依赖。
- 方案:将base环境替换为Python3.8版本,比如
AzureML-sklearn-0.24-ubuntu18.04-py38-cpu;或把environment.yml中的python=3.8改为python=3.7。
2. 确认condaDependenciesFile路径正确
RunConfig中condaDependenciesFile: environment.yml,但ScriptRunConfig的source_directory是/pipeline/task_scripts,如果environment.yml实际在/pipeline/task_runconfigs目录,需要调整路径为相对路径:
condaDependenciesFile: ../task_runconfigs/environment.yml
或者将environment.yml移动到task_scripts目录下。
3. 手动注册环境到Workspace
提前创建并注册环境,避免管道运行时环境创建出错:
from azureml.core import Environment # 从conda文件创建环境 env = Environment.from_conda_specification( name="phinmo_pipeline_env", file_path=os.path.join(WORKING_DIR, '/pipeline/task_runconfigs/environment.yml') ) # 注册到工作区 env.register(workspace=workspace) # 替换RunConfig中的环境 task_1_run_config.environment = env
4. 禁用环境缓存,强制重建
AzureML可能缓存旧环境,导致新依赖未安装。在PythonScriptStep中添加allow_reuse=False:
task_1_py_script_step = PythonScriptStep( name='Task_1_Step', script_name=task_1_script_run_config.script, source_directory=task_1_script_run_config.source_directory, compute_target=compute_target, allow_reuse=False )
或提交管道时强制重新生成输出:
pipeline_run = Experiment(workspace, 'Test_Run_New_Pipeline').submit( pipeline_run_config, regenerate_outputs=True )
5. 查看环境构建日志排查安装错误
在AzureML Studio中找到对应的运行任务,查看azureml-logs/20_image_build_log.txt,里面会记录环境创建的详细过程,确认pandas是否被正确安装,是否有安装失败的报错信息。
内容的提问来源于stack exchange,提问作者nils.hahn

