You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Shell脚本分批处理文件列表:每批4个的高效容错实现咨询

解决方案

原脚本核心问题分析

  1. 嵌套循环逻辑混乱:外层read逐行读取文件时,内层循环重复处理同一行数据,导致逻辑错误
  2. 破坏输入流:直接修改正在读取的list.txt(sed -i)会干扰read的读取行为,引发不可预测的结果
  3. 未等待后台任务完成:my_first_function "$files" &后台执行后未做等待,后续的日志移动、列表修改等操作可能在函数未完成时就执行,导致数据不一致
  4. 重复操作资源:find命令每次都移动所有.logs文件,会重复处理已移动的文件
  5. 不符合批次处理需求:my_second_function在单个文件处理后就执行,违背了“处理完4个文件后统一执行”的要求

优化实现方案

以下提供两种常用实现,分别适配串行/并行处理批次内文件的场景,兼顾执行效率与容错性:

方案1:串行处理批次内文件(稳定优先)

适合对文件处理顺序有要求,或单文件处理资源消耗较高的场景

#!/bin/bash

max_batch=4
Translated_files="/home/german_translated_files/"
list_file="list.txt"
# 创建临时文件存储处理失败的文件,用于后续重试
temp_list=$(mktemp)

# 脚本退出时自动清理临时文件
trap 'rm -f "$temp_list"' EXIT

while true; do
    # 从列表中读取最多4行作为当前批次
    mapfile -t batch_files -n "$max_batch" < "$list_file"
    # 批次为空则退出循环,所有文件处理完成
    [[ ${#batch_files[@]} -eq 0 ]] && break

    # 逐个处理批次内的文件
    for file in "${batch_files[@]}"; do
        # 执行翻译函数,检查执行结果
        if my_first_function "$file"; then
            echo "成功处理文件: $file"
        else
            echo "处理失败,将重试文件: $file" >&2
            # 失败文件写入临时列表,后续继续处理
            echo "$file" >> "$temp_list"
        fi
    done

    # 整批文件处理完成后,执行统一处理函数
    my_second_function

    # 移动日志文件,-n参数避免覆盖已存在文件
    find /home/german_translator/ -name '*.logs' -exec mv -n {} "$Translated_files" \;

    # 更新待处理文件列表
    if [[ -s "$temp_list" ]]; then
        # 将失败文件写回原列表,继续处理
        cat "$temp_list" > "$list_file"
        > "$temp_list" # 清空临时文件
    else
        # 无失败文件,清空原列表
        > "$list_file"
    fi
done

方案2:并行处理批次内文件(效率优先)

适合文件处理逻辑独立、资源占用较低的场景,可同时处理4个文件提升速度

#!/bin/bash

max_batch=4
Translated_files="/home/german_translated_files/"
list_file="list.txt"
temp_list=$(mktemp)

# 退出时清理临时文件+终止所有子进程
trap 'rm -f "$temp_list"; kill 0' EXIT

while true; do
    mapfile -t batch_files -n "$max_batch" < "$list_file"
    [[ ${#batch_files[@]} -eq 0 ]] && break

    # 并行启动当前批次的所有文件处理任务
    for file in "${batch_files[@]}"; do
        (
            if my_first_function "$file"; then
                echo "成功处理文件: $file"
            else
                echo "处理失败,将重试文件: $file" >&2
                echo "$file" >> "$temp_list"
            fi
        ) &
    done

    # 等待当前批次所有后台任务完成,再进入下一批
    wait

    # 执行统一处理函数
    my_second_function

    # 移动日志文件
    find /home/german_translator/ -name '*.logs' -exec mv -n {} "$Translated_files" \;

    # 更新待处理列表
    if [[ -s "$temp_list" ]]; then
        cat "$temp_list" > "$list_file"
        > "$temp_list"
    else
        > "$list_file"
    fi
done

关键注意事项

  1. 如果使用旧版本Bash(低于4.0),mapfile不可用,可替换为以下批量读取逻辑:
    batch_files=()
    count=0
    while IFS= read -r file && [[ $count -lt $max_batch ]]; do
        batch_files+=("$file")
        ((count++))
    done < "$list_file"
    
  2. 若无需保留原文件列表,可省略临时文件逻辑,直接处理后丢弃原列表
  3. mv -n参数可避免覆盖已存在文件,提升容错性;若需覆盖可移除该参数
  4. 可添加日志文件记录处理过程,例如将输出重定向到>> process.log 2>&1,方便后续排查问题

内容的提问来源于stack exchange,提问作者Paul C.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 16:48:24