如何用Python合并文本文件中每个ID行后的多行内容
用Python处理大型文本文件:合并ID行后的多行内容
针对你的需求,核心思路是逐行读取文件(避免一次性加载大文件占用过多内存),跟踪当前ID对应的待合并内容,遇到新的ID行时输出上一个ID的合并结果,具体实现如下:
代码实现
# 替换为你的输入文件路径 input_file = "input.txt" # 替换为你的输出文件路径 output_file = "output.txt" current_id = None merged_content = "" with open(input_file, 'r', encoding='utf-8') as infile, open(output_file, 'w', encoding='utf-8') as outfile: for line in infile: # 去掉行尾的换行符,保留内容本身 stripped_line = line.rstrip('\n') # 判断是否是ID行 if stripped_line.startswith("ID "): # 如果之前有未输出的合并内容,先写入文件 if current_id is not None: outfile.write(f"{merged_content}\n") # 写入当前ID行 outfile.write(f"{stripped_line}\n") # 重置变量,处理下一个ID current_id = stripped_line merged_content = "" else: # 非ID行,追加到合并内容中 merged_content += stripped_line # 处理最后一个ID的合并内容,避免遗漏 if current_id is not None: outfile.write(f"{merged_content}\n")
关键细节说明
- 逐行读取:使用
for line in infile的方式,每次只加载一行到内存,适合处理数千个ID的大型文件。 - 换行符处理:用
rstrip('\n')只去掉行尾的换行,保留内容中的其他空白(如果有的话),确保合并后的内容连贯。 - 边界处理:循环结束后额外处理最后一个ID的内容,避免文件末尾的合并内容被遗漏。
- 编码设置:指定
encoding='utf-8'确保兼容不同语言的文本内容,可根据实际文件编码调整。
内容的提问来源于stack exchange,提问作者Holly234
相关产品推荐
相关产品推荐

