You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python非正则文本解析性能优化:寻求更快实现方案

优化Python文本解析函数(无正则,提速)

核心优化思路

要同时把空格和逗号作为分隔符,且避免正则、提升速度,关键是减少冗余的字符串操作和中间变量。原实现中split()+join()的组合会产生额外的列表和字符串对象,开销较大。改用str.translate()一次性替换空白为逗号,再拆分过滤,能大幅提升效率。

优化后的代码

import string

# 预创建转换表:将所有空白字符(空格、换行、制表符等)映射为逗号
# 若只需处理空格,可改为 str.maketrans({' ': ','})
WHITESPACE_TO_COMMA = str.maketrans({c: ',' for c in string.whitespace})

def text_extractor(infile):
    output = []
    for line in infile:
        # 1. 替换所有空白为逗号 2. 按逗号拆分
        parts = line.translate(WHITESPACE_TO_COMMA).split(',')
        # 过滤空字符串,filter(None, ...)是高效的内置过滤操作
        filtered = list(filter(None, parts))
        if filtered:
            output.append(filtered)
    return output

def text_writer(input_data, outfile):
    for line in input_data:
        outfile.write(' '.join(line) + '\n')

# 调用示例
with open('./test.dat','rt') as infile, open('./test.out','wt') as outfile:
    output = text_extractor(infile)
    text_writer(output, outfile)

为什么更快?

  1. 单次遍历替换:str.translate()是底层C实现的字符串操作,仅需一次遍历就能完成所有空白到逗号的替换,比原方案中split()生成列表再join()拼接的两次遍历高效得多。
  2. 冗余操作移除:原方案中map(str.strip)是冗余的——替换空白为逗号后,拆分出的元素本身已无空格,无需额外strip。
  3. 高效过滤:filter(None, parts)利用内置过滤函数直接移除空字符串,比列表推导式的Python层循环更快。

输入输出验证

输入:

*HEADER, ./test/test.inp
1, 2, 3, 4,,,

a b c d e f ,

输出结构:

[["*HEADER", "./test/test.inp"],
 ["1","2","3","4"],
 ["a","b","c","d","e","f"]]

完全符合需求,且处理大文件时速度提升明显。


内容的提问来源于stack exchange,提问作者BurgerKing Lee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 15:41:09