如何复用requirements.txt指定的Python模块以避免重复下载提升构建速度?
复用已下载Python包以提升构建速度
问题场景
我有A、B两个文件夹,每个文件夹下包含若干子目录,且各自带有requirements.txt文件(比如A目录的文件包含boto3,B目录的包含pymongo==4.6.0)。每次运行bash脚本时,都会从互联网重新下载这些包,耗时很长。希望实现仅当requirements.txt中的包版本变更时才重新下载,否则复用已下载的包,以此提升构建速度。
相关文件示例:
cat /workspace/source/A/Lambda/requirements.txt boto3 cat /workspace/source/B/Lambda/requirements.txt pymongo==4.6.0
原bash脚本:
#!/bin/bash startdir=( "/workspace/source/A/Lambda" "/workspace/source/B/Lambda") SRC_DIR=/workspace/source/src_lambda if [ ! -d $SRC_DIR ] then mkdir $SRC_DIR else echo "Source Python Modules Directory Cleaning" cd "$SRC_DIR" && rm -rf *.zip cd .. fi for dir in "${startdir[@]}"; do ( cd "$dir" || continue for sub_dir in */ ; do cd "$sub_dir" || exit echo "---> Finding pip modules based on requirements.txt file" pip install --platform manylinux2014_x86_64 --implementation cp --python-version 3.11 --only-binary=:all: -r requirements.txt -t . zip -r "${sub_dir%/}.zip" . echo "---> Copying the zip files "${sub_dir%/}".zip into destination location "$SRC_DIR"" cp -r "${sub_dir%/}.zip" $SRC_DIR cd .. done ) done echo "zip files has been moved to SRC_DIR"
解决方案
核心思路是利用pip缓存复用已下载包,同时通过哈希校验判断requirements.txt是否变更,只有变更时才重新执行安装和打包操作。
具体实现步骤
- 指定pip持久缓存目录:避免pip默认缓存被清理,确保下载的包可以跨构建复用
- 生成
requirements.txt的哈希文件:每个子目录下生成requirements.hash,记录文件的哈希值,用于判断内容是否变更 - 修改脚本逻辑:先对比哈希值,只有当哈希变化时才执行
pip install和重新打包,否则直接复用已有的zip包(如果存在)
修改后的脚本
#!/bin/bash # 定义持久化pip缓存目录,可根据实际路径调整 PIP_CACHE_DIR="/workspace/pip_cache" mkdir -p "$PIP_CACHE_DIR" startdir=( "/workspace/source/A/Lambda" "/workspace/source/B/Lambda") SRC_DIR=/workspace/source/src_lambda if [ ! -d "$SRC_DIR" ]; then mkdir "$SRC_DIR" else echo "Source Python Modules Directory Cleaning" cd "$SRC_DIR" && rm -rf *.zip cd .. fi for dir in "${startdir[@]}"; do ( cd "$dir" || continue for sub_dir in */ ; do sub_dir_path=$(realpath "$sub_dir") cd "$sub_dir_path" || exit req_file="requirements.txt" hash_file="requirements.hash" zip_name="${sub_dir%/}.zip" dest_zip="$SRC_DIR/$zip_name" # 计算当前requirements.txt的哈希值 current_hash=$(sha256sum "$req_file" | awk '{print $1}') # 检查哈希文件是否存在,且哈希值是否一致,同时目标zip是否存在 if [ -f "$hash_file" ] && [ "$(cat "$hash_file")" = "$current_hash" ] && [ -f "$dest_zip" ]; then echo "---> Requirements未变更,复用已有的 $zip_name" cd .. continue fi # 若哈希变化或无缓存,执行安装 echo "---> Requirements已变更,重新安装依赖并打包" # 清理旧的依赖文件和zip包 rm -rf *.pyc __pycache__ *.zip pip install \ --platform manylinux2014_x86_64 \ --implementation cp \ --python-version 3.11 \ --only-binary=:all: \ --cache-dir "$PIP_CACHE_DIR" \ -r "$req_file" \ -t . # 生成新的哈希文件 echo "$current_hash" > "$hash_file" # 打包并复制到目标目录 zip -r "$zip_name" . echo "---> 复制 $zip_name 到 $SRC_DIR" cp "$zip_name" "$SRC_DIR" cd .. done ) done echo "所有zip包已同步到SRC_DIR"
关键说明
- pip缓存:通过
--cache-dir指定持久化目录,pip会自动复用该目录下已下载的包,无需重复从网络获取 - 哈希校验:使用
sha256sum生成requirements.txt的哈希值,只有当文件内容(包版本、新增/删除包)变化时,才会触发重新安装 - zip包复用:如果目标目录已经存在对应zip包且requirements未变更,直接跳过打包和复制步骤,进一步节省时间
内容的提问来源于stack exchange,提问作者Gowmi
相关产品推荐
相关产品推荐

