You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

通过Bash脚本执行cat合并大量文件是否更慢?求提速建议

问题解答

一、Bash脚本执行cat是否更慢或有局限?

直接在命令行和Bash脚本中执行cat本身性能差异可以忽略,你感觉到的速度差异大概率是以下原因导致:

  • 文件数量/参数限制:当前数据集的2000万个文件远超Bash的命令行参数长度上限(可通过getconf ARG_MAX查看,通常为几MB)。如果你的脚本里用了for f in this_files_*; do cat "$f" >> output.txt; done这类循环逐个处理文件,会频繁启动cat进程、反复打开/关闭输出文件,速度远慢于一次性批量处理;而上周的数据集文件数量少,命令行直接执行cat能一次性处理,自然更快。
  • 脚本调试模式:若脚本开启了set -x等调试选项,实时输出命令会额外消耗资源拖慢速度。
  • 环境差异:脚本的环境变量(如PATH)与命令行不同,可能调用了带额外逻辑的cat别名或不同版本工具。

二、更快的实现方案

1. Bash方案(适配大量文件+减少IO开销)

针对大量小文件,用find配合批量执行逻辑,避免参数过长问题,同时减少进程启动次数:

# 方案1:find + 批量exec
find . -maxdepth 1 -name "this_files_*" -type f -exec cat {} + > output.txt

# 方案2:find + xargs(处理特殊文件名更稳妥)
find . -maxdepth 1 -name "this_files_*" -type f -print0 | xargs -0 cat > output.txt

如果后续要排序,直接用管道跳过中间文件,节省磁盘IO时间:

find . -maxdepth 1 -name "this_files_*" -type f -exec cat {} + | sort > output_sorted.txt

2. Python方案(低内存+高效IO)

用Python流式读写,避免频繁启动外部进程,内存占用可控:

import os

output_path = "output.txt"
pattern_prefix = "this_files_"

# 流式写入输出文件,全程保持文件打开
with open(output_path, "wb") as out_f:
    for entry in os.scandir("."):
        if entry.is_file() and entry.name.startswith(pattern_prefix):
            with open(entry.path, "rb") as in_f:
                # 每次读取64KB块,可根据磁盘性能调整
                while chunk := in_f.read(65536):
                    out_f.write(chunk)

如需直接排序,结合subprocess调用系统sort(避免加载全量数据到内存):

import os
import subprocess

pattern_prefix = "this_files_"
# 启动sort进程,直接写入输出文件
sort_process = subprocess.Popen(
    ["sort", "-o", "output_sorted.txt"],
    stdin=subprocess.PIPE,
    bufsize=65536
)

with sort_process.stdin as stdin:
    for entry in os.scandir("."):
        if entry.is_file() and entry.name.startswith(pattern_prefix):
            with open(entry.path, "rb") as in_f:
                while chunk := in_f.read(65536):
                    stdin.write(chunk)

sort_process.wait()

三、额外优化建议

  • 确保使用SSD存储,机械硬盘在处理大量小文件时性能会严重受限。
  • 关闭系统中不必要的进程,降低CPU/磁盘负载。
  • 排序时可启用多线程加速:
    find ... -exec cat {} + | sort --parallel=8 > output_sorted.txt
    
    其中8替换为你的CPU核心数。

内容的提问来源于stack exchange,提问作者Ranger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 16:10:26