如何在Python3.6版AWS Lambda中导入指定库并解决包大小超限问题
The biggest pain point here is that these libraries rely on native C extensions—installing them on your local machine and uploading won't work, since they need to be compatible with Lambda's Amazon Linux 2 environment. Here are the most reliable ways to make this work:
Use AWS Lambda Layers (Recommended)
Layers let you package dependencies separately from your core code, making reuse easier and keeping your deployment package tiny. Here's how to build compatible layers:- Use a Docker container that matches Lambda's Python 3.6 environment to install dependencies. Run this in your terminal:
docker run -v "$PWD":/var/task public.ecr.aws/sam/build-python3.6:latest /bin/bash -c " mkdir -p layer/python pip install pandas numpy nltk spacy --target ./layer/python # Download Spacy's lightweight model (swap for a larger one if needed) python -m spacy download en_core_web_sm -d ./layer/python/lib/python3.6/site-packages # Grab only the NLTK corpora you actually use (e.g., punkt, stopwords) python -m nltk.downloader punkt stopwords -d ./layer/python/lib/python3.6/site-packages/nltk_data " - Zip the
layerdirectory (ensure thepythonfolder is at the root of the zip, not nested). - Upload this zip as a new Lambda layer in the AWS Console, then attach it to your Python 3.6 Lambda function.
- In your function code, import libraries normally:
import pandas as pd,import spacy, etc.
- Use a Docker container that matches Lambda's Python 3.6 environment to install dependencies. Run this in your terminal:
Optimized Direct Packaging
If you'd rather skip layers, build compatible packages locally using platform-specificpipflags:mkdir package pip install pandas numpy nltk spacy --target ./package --platform manylinux2014_x86_64 --only-binary=:all: --python-version 3.6 # Add Spacy model and NLTK data as shown above cp your_function_code.py package/ cd package && zip -r ../lambda_function.zip .Upload
lambda_function.zipto Lambda. Note this creates a larger package, so layers are better for long-term maintainability.
First, let's clarify Lambda's key limits to frame the solution:
- Maximum uploaded zip package size: 50MB
- Maximum unzipped size (including layers): 250MB
Your 62MB zip exceeds the upload limit, and the unzipped version is likely pushing close to or over the 250MB cap. Here are the best fixes:
Split Dependencies into Layers
Break your libraries into multiple focused layers (e.g., one for data processing: pandas, numpy, pytz; one for NLP: nltk, spacy; one for databases: psycopg2). Each layer is uploaded separately, so your main function zip will only contain your code (usually just a few KB). This bypasses the 50MB zip limit entirely, and layers count towards the 250MB unzipped total—you'll have plenty of headroom here.Trim Dependency Size
Cut down the bulk of your packages to fit within limits:- Use
pip install --no-depsto skip unnecessary transitive dependencies (only do this if you're certain you don't need them). - Delete redundant files: remove
__pycache__folders, test suites, documentation, and example directories. Runfind package -type d -name "__pycache__" -exec rm -r {} +manually or use tools likepyclean. - Swap to lightweight Spacy models (e.g.,
en_core_web_smis ~10MB, vs ~1.5GB foren_core_web_lg). - For NLTK, only download the specific corpora you use—avoid the full
allpackage.
- Use
Deploy via Container Image
Lambda supports container images up to 10GB, eliminating both the 50MB zip and 250MB unzipped limits. Here's a quick workflow:- Create a
Dockerfilebased on AWS's official Python 3.6 Lambda image:FROM public.ecr.aws/lambda/python:3.6 # Install dependencies COPY requirements.txt . RUN pip install -r requirements.txt --target "${LAMBDA_TASK_ROOT}" # Copy function code COPY your_function_code.py ${LAMBDA_TASK_ROOT} # Set the Lambda handler CMD ["your_function_code.lambda_handler"] - Build the image, push it to Amazon ECR, then deploy it as a Lambda function via the AWS Console. This is perfect for complex dependency stacks that can't be trimmed enough.
- Create a
Split Your Function into Multiple Lambdas
If your code can be logically divided (e.g., one function for data cleaning with pandas/numpy, another for NLP with spacy/nltk, a third for database work with psycopg2), create separate Lambdas each using only the dependencies they need. Use AWS Step Functions or API Gateway to orchestrate these functions in sequence. This keeps each individual deployment package well under size limits.
内容的提问来源于stack exchange,提问作者Ritu Chawla

