如何高效实现两组TXT文件的同名内容追加合并?
高效实现批量文本内容追加的Python方案
需求说明
我有两组.txt文件:
- 第一组文件按字母顺序排序,文件名覆盖从
aaa.txt到zzz.txt的范围 - 第二组文件存储在其他位置,包含
ant.txt、bat.txt、cat.txt等零散文件
需要完成的操作:将第二组中与第一组同名的文件内容,追加到第一组的对应文件中;第二组中第一组没有的文件直接忽略。
常规方案的问题
常规嵌套循环的实现逻辑效率极低:
for file in second_group: for file in first_group: # 检查文件名是否一致,一致则追加内容
这种方式的时间复杂度是O(m*n),完全没必要——就像人类找cat.txt不会从aaa.txt挨个翻,直接定位目标才是高效做法。
优化方案
方案1:直接文件名匹配(最优)
利用文件系统的快速查找能力,先把第一组的文件名存入集合(集合查找时间为O(1)),再遍历第二组文件直接判断是否存在匹配项,无需遍历第一组所有文件。
代码实现(用pathlib,简洁直观)
from pathlib import Path # 替换为实际的目录路径 first_group_dir = Path("/your/first/group/path") second_group_dir = Path("/your/second/group/path") # 获取第一组所有txt文件名的集合(仅保留文件名,不含路径) first_file_names = {f.name for f in first_group_dir.glob("*.txt")} # 遍历第二组的每个txt文件 for source_file in second_group_dir.glob("*.txt"): file_name = source_file.name if file_name in first_file_names: # 构造第一组对应文件的路径 target_file = first_group_dir / file_name # 追加内容:指定utf-8编码避免乱码 with open(target_file, 'a', encoding='utf-8') as target_f, open(source_file, 'r', encoding='utf-8') as source_f: target_f.write(source_f.read())
效率优势
整体时间复杂度为O(n)(n为第二组文件数量),远优于嵌套循环的O(m*n),且直接通过路径定位目标文件,完全避免无效遍历。
方案2:保留原第一组文件(对应你设想的第三目录)
如果不想修改原第一组文件,可以先将第一组文件复制到第三目录,再把第二组内容追加到第三目录的对应文件中,还可选择删除原第一组已处理的文件。
代码实现
from pathlib import Path import shutil # 替换为实际的目录路径 first_group_dir = Path("/your/first/group/path") second_group_dir = Path("/your/second/group/path") third_group_dir = Path("/your/third/group/path") # 创建第三目录(不存在则自动创建) third_group_dir.mkdir(exist_ok=True) # 先把第一组所有txt文件复制到第三目录 for file in first_group_dir.glob("*.txt"): shutil.copy(file, third_group_dir / file.name) # 遍历第二组文件,追加到第三目录对应文件 for source_file in second_group_dir.glob("*.txt"): file_name = source_file.name target_file = third_group_dir / file_name if target_file.exists(): with open(target_file, 'a', encoding='utf-8') as target_f, open(source_file, 'r', encoding='utf-8') as source_f: target_f.write(source_f.read()) # 可选:删除原第一组的对应文件 (first_group_dir / file_name).unlink(missing_ok=True)
额外提示
- 若第一组文件数量极大(数十万级),
pathlib.glob比os.listdir更稳定,还支持递归查找(如需) - 务必指定
encoding='utf-8',避免不同系统下的编码乱码问题 - 可添加异常处理增强脚本健壮性:
try: with open(target_file, 'a', encoding='utf-8') as target_f, open(source_file, 'r', encoding='utf-8') as source_f: target_f.write(source_f.read()) except Exception as e: print(f"处理文件 {file_name} 时出错: {e}")
内容的提问来源于stack exchange,提问作者Allure
相关产品推荐
相关产品推荐

