如何在Bash中预计算批量压缩文件最终大小并按预设大小分组归档
按文件大小分批压缩的Bash脚本方案
我有一个目录里约15GB的各类文件,需要压缩成单个大小不超过500MB的Zip归档文件。目前用的Bash脚本是按固定最大文件数分批压缩,虽然能正常运行,但因为单个文件大小差异大,生成的归档数量远超预期。想找一个能按预设最大归档大小计算每组文件数量的高效Bash脚本,或者能累计当前批次文件总大小,达到预设值就压缩的脚本,同时也想知道有没有最优解决方案。
当前使用的脚本
#!/bin/bash # Set the directory path to the folder containing the files to be compressed DIR_PATH="/path" # Set the name prefix of the output archive files ARCHIVE_PREFIX="archive" # Set the maximum number of files per batch MAX_FILES=1000 # Change directory to the specified path cd "$DIR_PATH" # Get a list of all files in the directory files=( * ) # Calculate the number of batches of files num_batches=$(( (${#files[@]} + $MAX_FILES - 1) / $MAX_FILES )) # Loop through each batch of files for (( i=0; i<$num_batches; i++ )); do # Set the start and end indices of the batch start=$(( $i * $MAX_FILES )) end=$(( ($i + 1) * $MAX_FILES - 1 )) # Check if the end index exceeds the number of files if (( $end >= ${#files[@]} )); then end=$(( ${#files[@]} - 1 )) fi # Create a compressed archive file for the batch of files archive_name="${ARCHIVE_PREFIX}_${i}.zip" tar -cvzf "$archive_name" "${files[@]:$start:$MAX_FILES}" done
注:原脚本用tar -cvzf生成的是tar.gz格式,不是标准Zip,这是一个小问题。
按文件大小累计的改进脚本
这个脚本会逐个遍历文件,累计当前批次的原始文件总大小,当达到预设阈值时就压缩该批次,同时处理单个文件超过阈值的情况:
#!/bin/bash # 目标文件目录 DIR_PATH="/path" # 输出归档文件的前缀 ARCHIVE_PREFIX="archive" # 单归档允许的最大原始文件总大小(500MB,单位字节) # 提示:如果文件多为高压缩率类型(如文本),可适当调大此值,压缩后会更小 MAX_TOTAL_SIZE=$((500 * 1024 * 1024)) # 切换到目标目录,失败则退出 cd "$DIR_PATH" || { echo "无法访问目录 $DIR_PATH"; exit 1; } # 初始化批次变量 current_batch_size=0 current_batch_files=() batch_number=0 # 遍历目录下所有文件(跳过子目录) for file in *; do [[ -f "$file" ]] || continue # 获取当前文件的大小(单位:字节) file_size=$(stat -c %s "$file") # 处理单个文件超过阈值的情况 if (( file_size > MAX_TOTAL_SIZE )); then echo "警告:文件 '$file' 大小超过单归档限制,将单独压缩" archive_name="${ARCHIVE_PREFIX}_${batch_number}_single.zip" zip -r "$archive_name" "$file" ((batch_number++)) continue fi # 如果加入当前文件后超过阈值,先压缩现有批次 if (( current_batch_size + file_size > MAX_TOTAL_SIZE )) && (( ${#current_batch_files[@]} > 0 )); then archive_name="${ARCHIVE_PREFIX}_${batch_number}.zip" zip -r "$archive_name" "${current_batch_files[@]}" # 重置批次状态 current_batch_size=0 current_batch_files=() ((batch_number++)) fi # 将当前文件加入批次 current_batch_files+=("$file") ((current_batch_size += file_size)) done # 处理最后一批剩余的文件 if (( ${#current_batch_files[@]} > 0 )); then archive_name="${ARCHIVE_PREFIX}_${batch_number}.zip" zip -r "$archive_name" "${current_batch_files[@]}" fi
脚本关键点说明
- 用
stat -c %s准确获取文件字节大小,避免文件名空格导致的错误 - 自动跳过子目录,只处理文件
- 单独处理超大文件,避免无法压缩的情况
- 使用
zip命令生成标准Zip归档,修正原脚本的格式问题 - 用数组存储当前批次文件,完美兼容带空格、特殊字符的文件名
最优方案建议
适配压缩率调整阈值:
不同类型文件压缩率差异极大(文本可达70%以上压缩率,视频/图片几乎无压缩)。可以先随机选几个文件测试压缩率,再调整MAX_TOTAL_SIZE——比如如果平均压缩率是50%,可以把阈值设为1GB,这样压缩后刚好接近500MB。提升压缩效率:
- 如果CPU有多个核心,推荐用
7z命令替代zip,支持多线程压缩且压缩率更高:
其中7z a -mx=5 -mmt=4 "$archive_name" "${current_batch_files[@]}"-mx=5是中等压缩级别(平衡速度和压缩率),-mmt=4是启用4线程。 - 若坚持用
zip,部分版本支持-T参数启用多线程,可尝试添加。
- 如果CPU有多个核心,推荐用
避免重复操作:
可以在压缩前检查是否已有同名归档文件,避免误覆盖,比如在生成archive_name后添加:if [[ -f "$archive_name" ]]; then echo "归档文件 $archive_name 已存在,跳过此批次" current_batch_size=0 current_batch_files=() ((batch_number++)) continue fi日志记录:
把压缩过程的输出重定向到日志文件,方便后续排查问题:zip -r "$archive_name" "${current_batch_files[@]}" >> compression_log.txt 2>&1
内容的提问来源于stack exchange,提问作者Ballyhoo
相关产品推荐
相关产品推荐

