如何为AWS Lambda制作轻量化SciPy包并解决numpy依赖报错
AWS Lambda精简scipy层相关问题解决方案
问题背景
AWS Lambda要求压缩后包大小不超过50MB,需使用体积较大的scipy包,因此精简scipy(仅保留scipy.sparse和scipy.optimize)制作自定义层,避免使用容器镜像(不熟悉),且无法切换至其他服务。
已执行操作
- 制作压缩包:在
aws-layer目录下执行pip3 install -r requirements.txt --target aws-layer/python/lib/python3.11/site-packages(requirements.txt仅含scipy),下载scipy和numpy后,删除scipy下除sparse、optimize外的所有文件夹,执行zip -r9 lambda-layer.zip .,最终包大小37MB,符合限制。 - 创建Lambda函数:运行时为python3.11,测试代码如下,添加层前出现预期的
No module named 'scipy'错误:
import json import scipy.sparse from scipy.optimize import milp, LinearConstraint, Bounds import numpy def lambda_handler(event, context): return { 'statusCode': 200, 'body': json.dumps('Hello from Lambda!') }
- 导入自定义层:创建运行时为python3.11的Lambda层,上传压缩包并关联到函数。
- 测试报错:层上传成功,但numpy导入失败,报错信息:
{ "errorMessage": "Unable to import module 'lambda_function': \n\nIMPORTANT: PLEASE READ THIS FOR ADVICE ON HOW TO SOLVE THIS ISSUE!\n\nImporting the numpy C-extensions failed. This error can happen for\nmany reasons, often due to issues with your setup or how NumPy was\ninstalled.\n\nWe have compiled some common reasons and troubleshooting tips at:\n\n https://numpy.org/devdocs/user/troubleshooting-importerror.html\n\nPlease note and check the following:\n\n * The Python version is: Python3.11 from \"/var/lang/bin/python3.11\"\n * The NumPy version is: \"1.21.6\"\n\nand make sure that they are the versions you expect.\nPlease carefully study the documentation linked above for further help.\n\nOriginal error was: No module named 'numpy.core._multiarray_umath'\n", "errorType": "Runtime.ImportModuleError", "requestId": "bde94bb7-ba45-4019-bf0f-81ef08c7124b", "stackTrace": [] }
- 尝试内置层:Lambda内置层
AWSSDKPandas-Python311可正常导入numpy,但即使移除自定义层中的numpy,同时添加该内置层和自定义scipy层时,总空间超出限制。
解决方案建议
1. 修复numpy导入错误
核心原因是本地安装的numpy与Lambda的python3.11环境不兼容,或打包时丢失了numpy的C扩展文件。解决步骤:
- 使用Lambda兼容的环境构建依赖:改用Amazon Linux 2容器镜像安装,确保编译的C扩展与Lambda运行时匹配:
# 拉取Amazon Linux 2镜像并挂载本地目录 docker run -v "$PWD/aws-layer:/aws-layer" -it amazonlinux:2 # 在容器内安装python3.11和pip yum install -y python3.11 python3.11-pip # 切换到目标目录安装依赖 cd /aws-layer python3.11 -m pip install -r requirements.txt --target python/lib/python3.11/site-packages # 精简scipy(删除不需要的子模块) rm -rf python/lib/python3.11/site-packages/scipy/{cluster,constants,fft,interpolate,io,linalg,ndimage,odr,signal,spatial,special,stats} # 打包层文件 zip -r9 lambda-layer.zip python/ - 检查numpy完整性:打包前确认
numpy/core/_multiarray_umath.cpython-311-x86_64-linux-gnu.so文件存在,切勿误删numpy核心文件。
2. 进一步精简scipy层
除删除scipy子模块外,可做以下优化:
- 删除scipy和numpy中的文档、测试文件:清理所有
*.md、*.rst、test_*.py文件,以及docs、tests文件夹; - 排除缓存文件:打包时用
zip -r9 --exclude="*__pycache__*" --exclude="*.pyc"命令,排除Python缓存文件(可选,Lambda支持pyc文件); - 指定轻量版本:安装scipy的旧版本(如1.10.x),新版本通常体积更大;
- 清理冗余依赖:检查requirements.txt,确保仅安装必要包,避免依赖传递带来的冗余。
3. 拆分Lambda函数的可行性
完全可行,架构设计如下:
- 函数A:部署精简后的scipy层,负责执行scipy相关计算逻辑;
- 函数B:使用内置pandas层,负责处理numpy/pandas相关逻辑;
- 函数C:作为协调器,接收请求后同步调用函数A和B,整合结果返回。
- 实现细节:在函数C中使用AWS SDK调用Lambda Invoke API触发A、B,数据量较小时用JSON传递,大数据则存储至S3后传递文件路径;添加错误重试机制,避免单次调用失败导致整体流程中断。
4. 容器镜像入门操作步骤
若以上方案均无法满足需求,可尝试容器镜像,步骤如下:
- 创建
Dockerfile:基于AWS官方Lambda Python镜像:FROM public.ecr.aws/lambda/python:3.11 # 复制依赖清单 COPY requirements.txt ${LAMBDA_TASK_ROOT} # 安装依赖 RUN pip3 install -r requirements.txt --target "${LAMBDA_TASK_ROOT}" # 复制Lambda函数代码 COPY lambda_function.py ${LAMBDA_TASK_ROOT} # 设置函数入口 CMD ["lambda_function.lambda_handler"] - 构建镜像:执行
docker build -t lambda-scipy .; - 推送镜像至AWS ECR:
- 在AWS控制台创建ECR私有仓库;
- 按照控制台提示的命令登录ECR,示例:
aws ecr get-login-password --region us-east-1 | docker login --username AWS --password-stdin 123456789012.dkr.ecr.us-east-1.amazonaws.com; - 给镜像打标签:
docker tag lambda-scipy:latest 123456789012.dkr.ecr.us-east-1.amazonaws.com/lambda-scipy:latest; - 推送镜像:
docker push 123456789012.dkr.ecr.us-east-1.amazonaws.com/lambda-scipy:latest;
- 创建Lambda函数:在控制台选择“容器镜像”选项,选择推送的ECR镜像完成创建。
内容的提问来源于stack exchange,提问作者CoolGuyHasChillDay
相关产品推荐
相关产品推荐

