You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

EMR PySpark集群模式下导入boto3失败的解决请求

问题分析
  1. Bootstrap脚本使用pip3 install -U boto3 botocore --user安装包时,默认会将包安装到执行用户的个人目录(若以root执行则是/root/.local/lib/pythonX.X/site-packages),但Spark集群模式下任务以hadoop用户运行,无法访问该目录下的包。
  2. 虽配置了PYSPARK_PYTHON=/usr/bin/python3,但该配置未正确覆盖所有节点的Spark环境,Executor端的Python环境未同步该路径。
  3. hadoop用户默认的pip3对应/bin/python3环境,而Bootstrap脚本安装的是/usr/bin/python3的环境,两者不属于同一Python环境。
解决方案

以下是无需修改simple.py的集群模式解决方法:

方法一:修改Bootstrap脚本全局安装Python包

将Bootstrap脚本改为全局安装(移除--user参数),确保所有用户(包括hadoop)都能访问boto3:

#!/bin/bash

echo -e 'Installing Boto3... \n'
which pip3
which python3
pip3 install -U boto3 botocore

若遇到权限问题,可添加sudo:

#!/bin/bash

echo -e 'Installing Boto3... \n'
which pip3
which python3
sudo pip3 install -U boto3 botocore

EMR 6.8.0默认Python3环境下,全局安装后所有节点的Python环境都会包含boto3,Spark集群模式运行时即可找到该包。

方法二:统一Python环境路径并确保Spark配置生效

  1. 修改Bootstrap脚本,将boto3安装到hadoop用户的Python环境,同时统一Python路径:
#!/bin/bash

# 切换到hadoop用户执行安装
su - hadoop -c "pip3 install -U boto3 botocore"
# 或指定/usr/bin/python3的pip安装到hadoop用户目录
/usr/bin/pip3 install -U boto3 botocore --user=hadoop
  1. 调整EMR的Spark配置,确保Master和Executor都使用相同的Python路径,在--configurations中补充如下配置:
[
  {
    "Classification": "spark-env",
    "Configurations": [
      {
        "Classification": "export",
        "Properties": {
          "PYSPARK_PYTHON": "/usr/bin/python3",
          "PYSPARK_DRIVER_PYTHON": "/usr/bin/python3"
        }
      }
    ]
  },
  {
    "Classification": "spark-defaults",
    "Properties": {
      "spark.executorEnv.PYSPARK_PYTHON": "/usr/bin/python3",
      "spark.driverEnv.PYSPARK_PYTHON": "/usr/bin/python3"
    }
  }
]

配置后,Spark的Driver和Executor都会使用/usr/bin/python3,而Bootstrap脚本也将boto3安装到该环境的hadoop用户目录下,集群模式即可找到包。

方法三:使用Spark的--archives参数打包Python环境

若不想修改Bootstrap脚本,可在提交任务时将包含boto3的Python环境打包成zip,通过--archives传递给集群:

  1. 在主节点以hadoop用户安装boto3到指定目录并打包:
su - hadoop
mkdir -p ./python_env
pip3 install boto3 botocore --target=./python_env
zip -r python_env.zip ./python_env
  1. 将zip包上传到S3,提交任务时添加参数:
spark-submit --deploy-mode cluster \
  --archives s3://your-bucket/python_env.zip#python_env \
  --conf spark.executorEnv.PYTHONPATH=./python_env \
  --conf spark.driverEnv.PYTHONPATH=./python_env \
  s3://something/py-spark/simple.py

此方法无需修改集群配置和Bootstrap脚本,仅在提交任务时传递环境,适合临时场景。

内容的提问来源于stack exchange,提问作者Flo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 00:15:22