You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按顺序遍历连续目录并处理其中的*.gz文件?

问题

当前处于目录a中,a目录下存在固定目录b,b目录下包含d091至d099的连续子目录,每个d目录下有不同的*.gz文件。需要从这些文件名中提取数据并按顺序输出,但原脚本处理的文件是无序的,输出的data_record内容如下:

ISTA00TUR_R_20190940000_01D_30S_MO.crx.gz
ISTA00TUR_R_20190990000_01D_30S_MO.crx.gz
ISTA00TUR_R_20190970000_01D_30S_MO.crx.gz
ISTA00TUR_R_20190920000_01D_30S_MO.crx.gz
ISTA00TUR_R_20190980000_01D_30S_MO.crx.gz
ISTA00TUR_R_20190910000_01D_30S_MO.crx.gz
ISTA00TUR_R_20190960000_01D_30S_MO.crx.gz
ISTA00TUR_R_20190930000_01D_30S_MO.crx.gz
ISTA00TUR_R_20190950000_01D_30S_MO.crx.gz

原脚本内容:

#!/bin/bash

for file in $(find . -name "*.gz"); do
  file=$(basename ${file})
       echo ${file:0:4} | tee -a receiver_ids > log
       echo ${file:16:17} | tee -a doy > log2
       echo ${file:0:100} | tee -a data_record > log3
done
cut -c 1-3 < doy > doy2
cut -c 1-23 < data_record > summary_name
解决方案

方法1:按已知目录顺序遍历(推荐)

既然已经明确b下是d091到d099的连续目录,直接按顺序遍历这些目录,就能保证文件处理的顺序,比用find再排序更高效:

#!/bin/bash

# 先清空所有输出文件,避免残留数据干扰结果
> receiver_ids
> doy
> data_record
> log
> log2
> log3

# 按d091到d099的顺序遍历目标目录
for dir in b/d09{1..9}; do
  # 遍历当前目录下的gz文件
  for file in "$dir"/*.gz; do
    # 跳过不存在的文件(防止部分目录无gz文件的情况)
    [[ -f "$file" ]] || continue
    base=$(basename "$file")
    # 写入数据并同步输出到对应log文件
    echo "${base:0:4}" | tee -a receiver_ids log
    echo "${base:16:17}" | tee -a doy log2
    echo "$base" | tee -a data_record log3
  done
done

# 执行后续的cut处理
cut -c 1-3 < doy > doy2
cut -c 1-23 < data_record > summary_name

方法2:对find结果排序后处理

如果一定要用find命令,可将find返回的结果通过sort排序后再处理:

#!/bin/bash

# 清空输出文件
> receiver_ids
> doy
> data_record
> log
> log2
> log3

# find结果排序后逐行处理,避免文件名含特殊字符的问题
find . -name "*.gz" | sort | while read -r file; do
  base=$(basename "$file")
  echo "${base:0:4}" | tee -a receiver_ids log
  echo "${base:16:17}" | tee -a doy log2
  echo "$base" | tee -a data_record log3
done

cut -c 1-3 < doy > doy2
cut -c 1-23 < data_record > summary_name

关键说明

  • 原脚本的核心问题是find返回的文件顺序由文件系统存储逻辑决定,无法保证有序,必须主动添加排序逻辑。
  • 方法1更贴合你的目录结构场景,完全按预期的d091到d099顺序处理,不会出现意外排序问题。
  • 脚本开头添加清空操作,避免之前运行的残留数据影响最终结果。
  • 用while read -r file代替for file in $(find ...),是更稳健的写法,可兼容文件名含空格、特殊字符的情况。

内容的提问来源于stack exchange,提问作者zeldazonk32

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 18:31:07