You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python大数据集正则替换内存溢出问题及循环改写求助

处理亿级代码行的函数名替换问题

需求

读取每一行代码,仅对以)结尾的行进行处理:将该行中**.或空格之后、(之前**的单词替换为function_call。

原方案的问题

原正则列表推导式处理百万级代码行正常,但处理1亿行时触发MemoryError——因为列表推导式会一次性把所有处理结果存入内存,亿级数据的内存占用远超系统承载能力。

原正则代码:

import re
data = ["int k = b.k(parcel)",
"int k = kon(parcel)",
"int a", 
"int bds",
"obtain.appendFrom(parcel, dataPosition2, readInt2)",
"obtain desFrom(package, dataPosition2, readInt2)",
"int abd(callme)",
"int.dbd(callyou)",
"int throw new UnsupportedOperationException(you)",
"int throw new.UnsupportedOperationException(me)"]

c_data = [
    re.sub(r"(\w+)\s*\(", "function_call(", i)
    for i in data
]

错误循环代码的问题点

你尝试的for循环存在两处关键错误:

  1. 条件判断逻辑错误:if " " or "." in i and i.endswith(")") 等价于 if True or ((".") in i and i.endswith(")")),所有行都会进入处理分支
  2. 替换逻辑错误:i.replace(f"{i}","function_call") 是将整行替换为function_call,完全没有定位到需要替换的目标单词

错误代码:

g = []
for i in data:
    if " " or "." in i and i.endswith(")"):
          g = i.replace(f"{i}","function_call")
          clean_data.append(g)

正确的逐行处理方案

核心思路:不一次性存储所有处理结果,逐行读取、处理、输出(或写入文件),避免内存溢出。以下提供两种实现方式:

方案1:优化正则逐行处理

预编译正则提升效率,逐行处理后直接输出/写入文件,不缓存所有结果:

import re

# 预编译正则,避免重复编译开销
pattern = re.compile(r"(\w+)\s*\(")

# 从文件逐行读取处理(推荐,适合亿级数据)
with open("input_code.txt", "r") as infile, open("output_code.txt", "w") as outfile:
    for line in infile:
        line = line.rstrip("\n")
        if line.endswith(")"):
            processed_line = pattern.sub("function_call(", line)
            outfile.write(processed_line + "\n")
        else:
            outfile.write(line + "\n")

# 若处理内存中的列表(注意:亿级列表本身会占内存,建议从文件读取)
for line in data:
    if line.endswith(")"):
        print(pattern.sub("function_call(", line))
    else:
        print(line)

方案2:手动字符串处理(无正则依赖)

通过字符串索引定位目标区域,适合对正则性能有顾虑的场景:

def process_single_line(line):
    if not line.endswith(")"):
        return line
    # 从后往前找第一个'('的位置
    open_paren_pos = line.rfind("(")
    if open_paren_pos == -1:
        return line
    # 找'('之前最后一个'.'或空格的位置
    split_pos = max(line.rfind(".", 0, open_paren_pos), line.rfind(" ", 0, open_paren_pos))
    # 跳过目标区域前的空白字符
    start_pos = split_pos + 1
    while start_pos < open_paren_pos and line[start_pos].isspace():
        start_pos += 1
    # 拼接处理后的字符串
    return line[:start_pos] + "function_call" + line[open_paren_pos:]

# 处理示例数据
data = ["int k = b.k(parcel)",
"int k = kon(parcel)",
"int a", 
"int bds",
"obtain.appendFrom(parcel, dataPosition2, readInt2)",
"obtain desFrom(package, dataPosition2, readInt2)",
"int abd(callme)",
"int.dbd(callyou)",
"int throw new UnsupportedOperationException(you)",
"int throw new.UnsupportedOperationException(me)"]

for line in data:
    print(process_single_line(line))

处理结果

两种方案均可输出符合预期的结果:

int k = b.function_call(parcel)
int k = function_call(parcel)
int a 
int bds
obtain.function_call(parcel, dataPosition2, readInt2)
obtain function_call(package, dataPosition2, readInt2)
int function_call(callme)
int.function_call(callyou)
int throw new function_call(you)
int throw new.function_call(me)

内容的提问来源于stack exchange,提问作者ya xi er

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 20:43:53