Python大数据集正则替换内存溢出问题及循环改写求助
处理亿级代码行的函数名替换问题
需求
读取每一行代码,仅对以)结尾的行进行处理:将该行中**.或空格之后、(之前**的单词替换为function_call。
原方案的问题
原正则列表推导式处理百万级代码行正常,但处理1亿行时触发MemoryError——因为列表推导式会一次性把所有处理结果存入内存,亿级数据的内存占用远超系统承载能力。
原正则代码:
import re data = ["int k = b.k(parcel)", "int k = kon(parcel)", "int a", "int bds", "obtain.appendFrom(parcel, dataPosition2, readInt2)", "obtain desFrom(package, dataPosition2, readInt2)", "int abd(callme)", "int.dbd(callyou)", "int throw new UnsupportedOperationException(you)", "int throw new.UnsupportedOperationException(me)"] c_data = [ re.sub(r"(\w+)\s*\(", "function_call(", i) for i in data ]
错误循环代码的问题点
你尝试的for循环存在两处关键错误:
- 条件判断逻辑错误:
if " " or "." in i and i.endswith(")")等价于if True or ((".") in i and i.endswith(")")),所有行都会进入处理分支 - 替换逻辑错误:
i.replace(f"{i}","function_call")是将整行替换为function_call,完全没有定位到需要替换的目标单词
错误代码:
g = [] for i in data: if " " or "." in i and i.endswith(")"): g = i.replace(f"{i}","function_call") clean_data.append(g)
正确的逐行处理方案
核心思路:不一次性存储所有处理结果,逐行读取、处理、输出(或写入文件),避免内存溢出。以下提供两种实现方式:
方案1:优化正则逐行处理
预编译正则提升效率,逐行处理后直接输出/写入文件,不缓存所有结果:
import re # 预编译正则,避免重复编译开销 pattern = re.compile(r"(\w+)\s*\(") # 从文件逐行读取处理(推荐,适合亿级数据) with open("input_code.txt", "r") as infile, open("output_code.txt", "w") as outfile: for line in infile: line = line.rstrip("\n") if line.endswith(")"): processed_line = pattern.sub("function_call(", line) outfile.write(processed_line + "\n") else: outfile.write(line + "\n") # 若处理内存中的列表(注意:亿级列表本身会占内存,建议从文件读取) for line in data: if line.endswith(")"): print(pattern.sub("function_call(", line)) else: print(line)
方案2:手动字符串处理(无正则依赖)
通过字符串索引定位目标区域,适合对正则性能有顾虑的场景:
def process_single_line(line): if not line.endswith(")"): return line # 从后往前找第一个'('的位置 open_paren_pos = line.rfind("(") if open_paren_pos == -1: return line # 找'('之前最后一个'.'或空格的位置 split_pos = max(line.rfind(".", 0, open_paren_pos), line.rfind(" ", 0, open_paren_pos)) # 跳过目标区域前的空白字符 start_pos = split_pos + 1 while start_pos < open_paren_pos and line[start_pos].isspace(): start_pos += 1 # 拼接处理后的字符串 return line[:start_pos] + "function_call" + line[open_paren_pos:] # 处理示例数据 data = ["int k = b.k(parcel)", "int k = kon(parcel)", "int a", "int bds", "obtain.appendFrom(parcel, dataPosition2, readInt2)", "obtain desFrom(package, dataPosition2, readInt2)", "int abd(callme)", "int.dbd(callyou)", "int throw new UnsupportedOperationException(you)", "int throw new.UnsupportedOperationException(me)"] for line in data: print(process_single_line(line))
处理结果
两种方案均可输出符合预期的结果:
int k = b.function_call(parcel) int k = function_call(parcel) int a int bds obtain.function_call(parcel, dataPosition2, readInt2) obtain function_call(package, dataPosition2, readInt2) int function_call(callme) int.function_call(callyou) int throw new function_call(you) int throw new.function_call(me)
内容的提问来源于stack exchange,提问作者ya xi er
相关产品推荐
相关产品推荐

