大XML文件奇数行trace字段替换/添加及Python代码优化求助
处理大型XML文件的行编辑优化方案
问题背景
有一个800MB以上的大型XML文本文件,数据格式示例如下:
<iden/><provider></provider><trace>065110d4-cec5-d433772ed57a</trace> <ServiceRQ>Some xml data</ServiceRQ> <iden/><provider></provider> <ServiceRQ>Some xml data</ServiceRQ>
需求:检查文件的偶数索引行(以0为起始索引,即第1、3、5...行):
- 若存在
<trace>标签,将标签内的内容替换为xyz - 若不存在
<trace>标签,在该行末尾添加<trace>xyz</trace>
原代码使用readlines()一次性读取整个文件,导致内存不足无法运行,需优化。
原问题代码
with open("Sample_xml.txt", 'r') as fp: output = fp.readlines() type(output) s = len(output) - 1 tc = 0 rq = 1 while (tc <= s) and (rq <= s): if tc % 2 == 0: a = (output[tc]) if a.find("<trace") != -1: a = re.sub('(?<=<trace>)(.*?)(?=</trace>)','xyz', a) print(a) elif a.find("<trace>") == -1: a = a.rstrip() + '<trace>xyz</trace>' +'\n' print(a) if rq % 2 != 0: b = (output[rq]) print(b) with open("Fin_xml.txt", "a") as myfile: myfile.write(a) myfile.write(b + '\n') tc += 2 rq += 2
优化后的代码
import re # 预编译正则表达式,提升重复匹配效率 trace_regex = re.compile(r'(?<=<trace>)(.*?)(?=</trace>)') # 同时打开输入和输出文件,逐行处理 with open("Sample_xml.txt", 'r', encoding='utf-8') as infile, \ open("Fin_xml.txt", 'w', encoding='utf-8') as outfile: for line_idx, line in enumerate(infile): if line_idx % 2 == 0: # 处理目标行:先去除末尾换行,避免重复添加 cleaned_line = line.rstrip('\n') if '<trace>' in cleaned_line: # 替换trace标签内的内容 modified_line = trace_regex.sub('xyz', cleaned_line) + '\n' else: # 添加trace标签到行尾 modified_line = f"{cleaned_line}<trace>xyz</trace>\n" outfile.write(modified_line) else: # 非目标行直接原样写入 outfile.write(line)
优化要点
- 逐行迭代读取:利用文件对象的迭代器特性,每次仅加载一行到内存,彻底解决大文件内存溢出问题。
- 预编译正则表达式:将需要重复使用的正则提前编译,避免每次处理行时重复编译,提升处理速度。
- 减少IO操作:原代码每次循环都打开/关闭输出文件,优化后一次性打开输出文件,大幅降低IO开销。
- 简化逻辑结构:使用
enumerate直接获取行索引,替代原代码的双计数器逻辑,代码更简洁易维护。 - 统一换行处理:明确处理行尾换行符,避免出现多余换行或缺失换行的问题。
内容的提问来源于stack exchange,提问作者Akash
相关产品推荐
相关产品推荐

