在AWS Lambda中使用Llama Index遭遇依赖与层大小问题
在AWS Lambda上部署Llama Index机器人的依赖层问题解决指南
问题概述
使用Llama Index在AWS Lambda搭建自定义机器人时,打包上传依赖层频繁遭遇「包未找到」错误;尝试替换为AWS官方提供的Pandas、NumPy层后,又触发「层大小超限」问题,反复调整后仍存在依赖缺失情况。
具体错误信息
NumPy导入错误
{ "errorMessage": "Unable to import module 'lambda_handler': Unable to import required dependencies:\nnumpy: \n\nIMPORTANT: PLEASE READ THIS FOR ADVICE ON HOW TO SOLVE THIS ISSUE!\n\nImporting the numpy C-extensions failed. This error can happen for\nmany reasons, often due to issues with your setup or how NumPy was\ninstalled.\n\nWe have compiled some common reasons and troubleshooting tips at:\n\n https://numpy.org/devdocs/user/troubleshooting-importerror.html\n\nPlease note and check the following:\n\n * The Python version is: Python3.8 from \"/var/lang/bin/python3.8\"\n * The NumPy version is: \"1.24.2\"\n\nand make sure that they are the versions you expect.\nPlease carefully study the documentation linked above for further help.\n\nOriginal error was: No module named 'numpy.core._multiarray_umath'\n", "errorType": "Runtime.ImportModuleError", "stackTrace": [] }
后续触发的其他错误
No module named 'pandas._libs.interval'No module named numexpr
示例Lambda代码
import json from llama_index import GPTSimpleVectorIndex import os import boto3 def lambda_handler(event, context): s3 = boto3.resource('s3') try: event = json.loads(event['body']) prompt = event.get("prompt", None) if(prompt is None): raise Exception("prompt is None") bucket_name = '' object_key = '' index_object = s3.Object(bucket_name, object_key) index_content = index_object.get()['Body'].read() index = GPTSimpleVectorIndex.load_from_disk(index_content) # Query the index RESPONSE = index.query(prompt) return { "headers": { "Access-Control-Allow-Headers": "Content-Type", "Access-Control-Allow-Origin": "*", "Access-Control-Allow-Methods": "POST" }, "statusCode": 200, "body": json.dumps( {'message': json.dumps({"message": RESPONSE})} ) } except Exception as e: print(e) return { "headers": { "Access-Control-Allow-Headers": "Content-Type", "Access-Control-Allow-Origin": "*", "Access-Control-Allow-Methods": "POST" }, "statusCode": 500, "body": json.dumps( {'message': json.dumps(e, default=str)} ) }
已尝试方案
- 删除Llama Index包内部分依赖,改用AWS官方Pandas、NumPy层
- 添加AWS NumPy层解决初始NumPy导入错误
- 添加AWSPandasSDK层解决Pandas错误,但触发层大小超限
- 移除独立NumPy层,保留AWSPandasSDK层,出现numexpr依赖缺失
解决方案
1. 解决跨平台编译导致的依赖错误
WSL2 Ubuntu环境打包的依赖与Lambda的Amazon Linux环境不兼容,会引发NumPy等带C扩展库的导入失败。需在Lambda兼容环境下打包:
- 拉取Lambda官方Python镜像:
docker run -v "$PWD":/var/task public.ecr.aws/lambda/python:3.8 bash - 在容器内创建虚拟环境并安装依赖:
python -m venv venv source venv/bin/activate pip install llama-index --target python/lib/python3.8/site-packages/ --no-cache-dir - 打包
python目录为zip文件,上传作为自定义Lambda层。
2. 优化依赖层大小,避免超限
- 清理冗余文件:
rm -rf python/lib/python3.8/site-packages/*/__pycache__ rm -rf python/lib/python3.8/site-packages/*/*.pyc - 仅安装核心依赖:若Llama Index的可选依赖非必需,可手动指定安装核心包:
pip install llama-index openai numpy pandas numexpr --target python/lib/python3.8/site-packages/ --no-cache-dir - 拆分层:将Llama Index核心依赖与大体积第三方库(如Pandas、NumPy)拆分为两个独立层,分别上传。
3. 修复numexpr依赖缺失问题
在容器打包时手动安装numexpr:
pip install numexpr --target python/lib/python3.8/site-packages/ --no-cache-dir
确保与Llama Index依赖一起打包上传。
4. Lambda代码逻辑修正
原代码中load_from_disk接收的是文件路径而非字节流,需先将S3读取的内容写入临时文件再加载:
import tempfile # ... 其他代码 index_content = index_object.get()['Body'].read() with tempfile.NamedTemporaryFile(mode='wb', delete=False) as tmp: tmp.write(index_content) tmp_path = tmp.name index = GPTSimpleVectorIndex.load_from_disk(tmp_path) # 清理临时文件 os.unlink(tmp_path)
内容的提问来源于stack exchange,提问作者Ishaan Tiwari
相关产品推荐
相关产品推荐

