You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Bash中预计算批量压缩文件最终大小并按预设大小分组归档

按文件大小分批压缩的Bash脚本方案

我有一个目录里约15GB的各类文件,需要压缩成单个大小不超过500MB的Zip归档文件。目前用的Bash脚本是按固定最大文件数分批压缩,虽然能正常运行,但因为单个文件大小差异大,生成的归档数量远超预期。想找一个能按预设最大归档大小计算每组文件数量的高效Bash脚本,或者能累计当前批次文件总大小,达到预设值就压缩的脚本,同时也想知道有没有最优解决方案。

当前使用的脚本

#!/bin/bash

# Set the directory path to the folder containing the files to be compressed
DIR_PATH="/path"

# Set the name prefix of the output archive files
ARCHIVE_PREFIX="archive"

# Set the maximum number of files per batch
MAX_FILES=1000

# Change directory to the specified path
cd "$DIR_PATH"

# Get a list of all files in the directory
files=( * )

# Calculate the number of batches of files
num_batches=$(( (${#files[@]} + $MAX_FILES - 1) / $MAX_FILES ))

# Loop through each batch of files
for (( i=0; i<$num_batches; i++ )); do
    # Set the start and end indices of the batch
    start=$(( $i * $MAX_FILES ))
    end=$(( ($i + 1) * $MAX_FILES - 1 ))
    
    # Check if the end index exceeds the number of files
    if (( $end >= ${#files[@]} )); then
        end=$(( ${#files[@]} - 1 ))
    fi
    
    # Create a compressed archive file for the batch of files
    archive_name="${ARCHIVE_PREFIX}_${i}.zip"
    tar -cvzf "$archive_name" "${files[@]:$start:$MAX_FILES}"
done

注:原脚本用tar -cvzf生成的是tar.gz格式,不是标准Zip,这是一个小问题。


按文件大小累计的改进脚本

这个脚本会逐个遍历文件,累计当前批次的原始文件总大小,当达到预设阈值时就压缩该批次,同时处理单个文件超过阈值的情况:

#!/bin/bash

# 目标文件目录
DIR_PATH="/path"
# 输出归档文件的前缀
ARCHIVE_PREFIX="archive"
# 单归档允许的最大原始文件总大小(500MB,单位字节)
# 提示:如果文件多为高压缩率类型(如文本),可适当调大此值,压缩后会更小
MAX_TOTAL_SIZE=$((500 * 1024 * 1024))

# 切换到目标目录,失败则退出
cd "$DIR_PATH" || { echo "无法访问目录 $DIR_PATH"; exit 1; }

# 初始化批次变量
current_batch_size=0
current_batch_files=()
batch_number=0

# 遍历目录下所有文件(跳过子目录)
for file in *; do
    [[ -f "$file" ]] || continue

    # 获取当前文件的大小(单位:字节)
    file_size=$(stat -c %s "$file")

    # 处理单个文件超过阈值的情况
    if (( file_size > MAX_TOTAL_SIZE )); then
        echo "警告:文件 '$file' 大小超过单归档限制,将单独压缩"
        archive_name="${ARCHIVE_PREFIX}_${batch_number}_single.zip"
        zip -r "$archive_name" "$file"
        ((batch_number++))
        continue
    fi

    # 如果加入当前文件后超过阈值,先压缩现有批次
    if (( current_batch_size + file_size > MAX_TOTAL_SIZE )) && (( ${#current_batch_files[@]} > 0 )); then
        archive_name="${ARCHIVE_PREFIX}_${batch_number}.zip"
        zip -r "$archive_name" "${current_batch_files[@]}"
        # 重置批次状态
        current_batch_size=0
        current_batch_files=()
        ((batch_number++))
    fi

    # 将当前文件加入批次
    current_batch_files+=("$file")
    ((current_batch_size += file_size))
done

# 处理最后一批剩余的文件
if (( ${#current_batch_files[@]} > 0 )); then
    archive_name="${ARCHIVE_PREFIX}_${batch_number}.zip"
    zip -r "$archive_name" "${current_batch_files[@]}"
fi

脚本关键点说明

  • 用stat -c %s准确获取文件字节大小,避免文件名空格导致的错误
  • 自动跳过子目录,只处理文件
  • 单独处理超大文件,避免无法压缩的情况
  • 使用zip命令生成标准Zip归档,修正原脚本的格式问题
  • 用数组存储当前批次文件,完美兼容带空格、特殊字符的文件名

最优方案建议

  1. 适配压缩率调整阈值:
    不同类型文件压缩率差异极大(文本可达70%以上压缩率,视频/图片几乎无压缩)。可以先随机选几个文件测试压缩率,再调整MAX_TOTAL_SIZE——比如如果平均压缩率是50%,可以把阈值设为1GB,这样压缩后刚好接近500MB。

  2. 提升压缩效率:

    • 如果CPU有多个核心,推荐用7z命令替代zip,支持多线程压缩且压缩率更高:
      7z a -mx=5 -mmt=4 "$archive_name" "${current_batch_files[@]}"
      
      其中-mx=5是中等压缩级别(平衡速度和压缩率),-mmt=4是启用4线程。
    • 若坚持用zip,部分版本支持-T参数启用多线程,可尝试添加。
  3. 避免重复操作:
    可以在压缩前检查是否已有同名归档文件,避免误覆盖,比如在生成archive_name后添加:

    if [[ -f "$archive_name" ]]; then
        echo "归档文件 $archive_name 已存在,跳过此批次"
        current_batch_size=0
        current_batch_files=()
        ((batch_number++))
        continue
    fi
    
  4. 日志记录:
    把压缩过程的输出重定向到日志文件,方便后续排查问题:

    zip -r "$archive_name" "${current_batch_files[@]}" >> compression_log.txt 2>&1
    

内容的提问来源于stack exchange,提问作者Ballyhoo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 06:37:01