You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效实现两组TXT文件的同名内容追加合并?

高效实现批量文本内容追加的Python方案

需求说明

我有两组.txt文件:

  • 第一组文件按字母顺序排序,文件名覆盖从aaa.txt到zzz.txt的范围
  • 第二组文件存储在其他位置,包含ant.txt、bat.txt、cat.txt等零散文件

需要完成的操作:将第二组中与第一组同名的文件内容,追加到第一组的对应文件中;第二组中第一组没有的文件直接忽略。

常规方案的问题

常规嵌套循环的实现逻辑效率极低:

for file in second_group:
   for file in first_group:
      # 检查文件名是否一致,一致则追加内容

这种方式的时间复杂度是O(m*n),完全没必要——就像人类找cat.txt不会从aaa.txt挨个翻,直接定位目标才是高效做法。

优化方案

方案1:直接文件名匹配(最优)

利用文件系统的快速查找能力,先把第一组的文件名存入集合(集合查找时间为O(1)),再遍历第二组文件直接判断是否存在匹配项,无需遍历第一组所有文件。

代码实现(用pathlib,简洁直观)

from pathlib import Path

# 替换为实际的目录路径
first_group_dir = Path("/your/first/group/path")
second_group_dir = Path("/your/second/group/path")

# 获取第一组所有txt文件名的集合(仅保留文件名,不含路径)
first_file_names = {f.name for f in first_group_dir.glob("*.txt")}

# 遍历第二组的每个txt文件
for source_file in second_group_dir.glob("*.txt"):
    file_name = source_file.name
    if file_name in first_file_names:
        # 构造第一组对应文件的路径
        target_file = first_group_dir / file_name
        # 追加内容:指定utf-8编码避免乱码
        with open(target_file, 'a', encoding='utf-8') as target_f, open(source_file, 'r', encoding='utf-8') as source_f:
            target_f.write(source_f.read())

效率优势

整体时间复杂度为O(n)(n为第二组文件数量),远优于嵌套循环的O(m*n),且直接通过路径定位目标文件,完全避免无效遍历。

方案2:保留原第一组文件(对应你设想的第三目录)

如果不想修改原第一组文件,可以先将第一组文件复制到第三目录,再把第二组内容追加到第三目录的对应文件中,还可选择删除原第一组已处理的文件。

代码实现

from pathlib import Path
import shutil

# 替换为实际的目录路径
first_group_dir = Path("/your/first/group/path")
second_group_dir = Path("/your/second/group/path")
third_group_dir = Path("/your/third/group/path")

# 创建第三目录(不存在则自动创建)
third_group_dir.mkdir(exist_ok=True)

# 先把第一组所有txt文件复制到第三目录
for file in first_group_dir.glob("*.txt"):
    shutil.copy(file, third_group_dir / file.name)

# 遍历第二组文件,追加到第三目录对应文件
for source_file in second_group_dir.glob("*.txt"):
    file_name = source_file.name
    target_file = third_group_dir / file_name
    if target_file.exists():
        with open(target_file, 'a', encoding='utf-8') as target_f, open(source_file, 'r', encoding='utf-8') as source_f:
            target_f.write(source_f.read())
        # 可选:删除原第一组的对应文件
        (first_group_dir / file_name).unlink(missing_ok=True)

额外提示

  • 若第一组文件数量极大(数十万级),pathlib.glob比os.listdir更稳定,还支持递归查找(如需)
  • 务必指定encoding='utf-8',避免不同系统下的编码乱码问题
  • 可添加异常处理增强脚本健壮性:
try:
    with open(target_file, 'a', encoding='utf-8') as target_f, open(source_file, 'r', encoding='utf-8') as source_f:
        target_f.write(source_f.read())
except Exception as e:
    print(f"处理文件 {file_name} 时出错: {e}")

内容的提问来源于stack exchange,提问作者Allure

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 23:55:16