Shell脚本分批处理文件列表:每批4个的高效容错实现咨询
解决方案
原脚本核心问题分析
- 嵌套循环逻辑混乱:外层
read逐行读取文件时,内层循环重复处理同一行数据,导致逻辑错误 - 破坏输入流:直接修改正在读取的
list.txt(sed -i)会干扰read的读取行为,引发不可预测的结果 - 未等待后台任务完成:
my_first_function "$files" &后台执行后未做等待,后续的日志移动、列表修改等操作可能在函数未完成时就执行,导致数据不一致 - 重复操作资源:
find命令每次都移动所有.logs文件,会重复处理已移动的文件 - 不符合批次处理需求:
my_second_function在单个文件处理后就执行,违背了“处理完4个文件后统一执行”的要求
优化实现方案
以下提供两种常用实现,分别适配串行/并行处理批次内文件的场景,兼顾执行效率与容错性:
方案1:串行处理批次内文件(稳定优先)
适合对文件处理顺序有要求,或单文件处理资源消耗较高的场景
#!/bin/bash max_batch=4 Translated_files="/home/german_translated_files/" list_file="list.txt" # 创建临时文件存储处理失败的文件,用于后续重试 temp_list=$(mktemp) # 脚本退出时自动清理临时文件 trap 'rm -f "$temp_list"' EXIT while true; do # 从列表中读取最多4行作为当前批次 mapfile -t batch_files -n "$max_batch" < "$list_file" # 批次为空则退出循环,所有文件处理完成 [[ ${#batch_files[@]} -eq 0 ]] && break # 逐个处理批次内的文件 for file in "${batch_files[@]}"; do # 执行翻译函数,检查执行结果 if my_first_function "$file"; then echo "成功处理文件: $file" else echo "处理失败,将重试文件: $file" >&2 # 失败文件写入临时列表,后续继续处理 echo "$file" >> "$temp_list" fi done # 整批文件处理完成后,执行统一处理函数 my_second_function # 移动日志文件,-n参数避免覆盖已存在文件 find /home/german_translator/ -name '*.logs' -exec mv -n {} "$Translated_files" \; # 更新待处理文件列表 if [[ -s "$temp_list" ]]; then # 将失败文件写回原列表,继续处理 cat "$temp_list" > "$list_file" > "$temp_list" # 清空临时文件 else # 无失败文件,清空原列表 > "$list_file" fi done
方案2:并行处理批次内文件(效率优先)
适合文件处理逻辑独立、资源占用较低的场景,可同时处理4个文件提升速度
#!/bin/bash max_batch=4 Translated_files="/home/german_translated_files/" list_file="list.txt" temp_list=$(mktemp) # 退出时清理临时文件+终止所有子进程 trap 'rm -f "$temp_list"; kill 0' EXIT while true; do mapfile -t batch_files -n "$max_batch" < "$list_file" [[ ${#batch_files[@]} -eq 0 ]] && break # 并行启动当前批次的所有文件处理任务 for file in "${batch_files[@]}"; do ( if my_first_function "$file"; then echo "成功处理文件: $file" else echo "处理失败,将重试文件: $file" >&2 echo "$file" >> "$temp_list" fi ) & done # 等待当前批次所有后台任务完成,再进入下一批 wait # 执行统一处理函数 my_second_function # 移动日志文件 find /home/german_translator/ -name '*.logs' -exec mv -n {} "$Translated_files" \; # 更新待处理列表 if [[ -s "$temp_list" ]]; then cat "$temp_list" > "$list_file" > "$temp_list" else > "$list_file" fi done
关键注意事项
- 如果使用旧版本Bash(低于4.0),
mapfile不可用,可替换为以下批量读取逻辑:batch_files=() count=0 while IFS= read -r file && [[ $count -lt $max_batch ]]; do batch_files+=("$file") ((count++)) done < "$list_file" - 若无需保留原文件列表,可省略临时文件逻辑,直接处理后丢弃原列表
mv -n参数可避免覆盖已存在文件,提升容错性;若需覆盖可移除该参数- 可添加日志文件记录处理过程,例如将输出重定向到
>> process.log 2>&1,方便后续排查问题
内容的提问来源于stack exchange,提问作者Paul C.
相关产品推荐
相关产品推荐

