You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python合并文本文件中每个ID行后的多行内容

用Python处理大型文本文件:合并ID行后的多行内容

针对你的需求,核心思路是逐行读取文件(避免一次性加载大文件占用过多内存),跟踪当前ID对应的待合并内容,遇到新的ID行时输出上一个ID的合并结果,具体实现如下:

代码实现

# 替换为你的输入文件路径
input_file = "input.txt"
# 替换为你的输出文件路径
output_file = "output.txt"

current_id = None
merged_content = ""

with open(input_file, 'r', encoding='utf-8') as infile, open(output_file, 'w', encoding='utf-8') as outfile:
    for line in infile:
        # 去掉行尾的换行符,保留内容本身
        stripped_line = line.rstrip('\n')
        # 判断是否是ID行
        if stripped_line.startswith("ID "):
            # 如果之前有未输出的合并内容,先写入文件
            if current_id is not None:
                outfile.write(f"{merged_content}\n")
            # 写入当前ID行
            outfile.write(f"{stripped_line}\n")
            # 重置变量,处理下一个ID
            current_id = stripped_line
            merged_content = ""
        else:
            # 非ID行,追加到合并内容中
            merged_content += stripped_line
    # 处理最后一个ID的合并内容,避免遗漏
    if current_id is not None:
        outfile.write(f"{merged_content}\n")

关键细节说明

  • 逐行读取:使用for line in infile的方式,每次只加载一行到内存,适合处理数千个ID的大型文件。
  • 换行符处理:用rstrip('\n')只去掉行尾的换行,保留内容中的其他空白(如果有的话),确保合并后的内容连贯。
  • 边界处理:循环结束后额外处理最后一个ID的内容,避免文件末尾的合并内容被遗漏。
  • 编码设置:指定encoding='utf-8'确保兼容不同语言的文本内容,可根据实际文件编码调整。

内容的提问来源于stack exchange,提问作者Holly234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 06:40:28