You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于文本文件与正则表达式创建DataFrame:Python代码优化咨询

大文本文件提取数据生成DataFrame的优化方案

问题背景

需从单份约500MB的文本文件(每行格式如下)中提取指定字段,生成pandas DataFrame:

2022-07-01,08:00:57.853, +12-34 = 1.11 (0. AA), a=0, b=1 cct= p=0 f=0 r=0 pb=0 pbb=0 prr=2569 du=89

需提取的字段与对应列名:

  • +12-34 → id
  • 1.11 → res
  • a=0 → a
  • b=1 → b
  • p=0 → p
  • f=0 → f
  • r=0 → r
  • pb=0 → pb
  • pbb=0 → pbb
  • prr=2569 → prr
  • du=89 → du

原代码存在内存占用高、处理速度慢的问题,以下是针对性优化建议:


核心优化方向

  • 预编译正则:避免重复编译正则表达式,减少不必要的计算开销
  • 逐行处理文件:不一次性加载所有行到内存,降低内存占用
  • 批量生成DataFrame:先收集所有数据到列表,最后一次性转换,替代循环拼接DataFrame
  • 简化提取逻辑:用单正则匹配整行关键内容,减少多次字符串拆分的耗时

优化后的代码

import glob
import re
import pandas as pd

# 预编译正则,匹配整行中需要提取的所有字段
line_pattern = re.compile(r'\+(\S+) = (\d\.\d{2}).*?a=(\w+),?\s*b=(\w+)\s*cct=\s*p=(\w+)\s*f=(\w+)\s*r=(\w+)\s*pb=(\w+)\s*pbb=(\w+)\s*prr=(\w+)\s*du=(\w+)')

path_torawfolder = "./your_raw_folder/"  # 替换为实际文件夹路径
files = glob.glob(path_torawfolder + "*.txt")

# 用列表批量存储所有数据字典
data_records = []

for fle in files:
    with open(fle, 'r', encoding='utf-8') as f:
        for line in f:
            line = line.strip()
            if not line:
                continue  # 跳过空行
            match_result = line_pattern.search(line)
            if match_result:
                # 按匹配组顺序映射到对应列名
                record = {
                    'id': '+' + match_result.group(1),
                    'res': match_result.group(2),
                    'a': match_result.group(3),
                    'b': match_result.group(4),
                    'p': match_result.group(5),
                    'f': match_result.group(6),
                    'r': match_result.group(7),
                    'pb': match_result.group(8),
                    'pbb': match_result.group(9),
                    'prr': match_result.group(10),
                    'du': match_result.group(11)
                }
                data_records.append(record)

# 一次性生成DataFrame,并指定数据类型减少内存占用
df = pd.DataFrame(data_records, dtype={
    'id': str,
    'res': float,
    'a': int,
    'b': int,
    'p': int,
    'f': int,
    'r': int,
    'pb': int,
    'pbb': int,
    'prr': int,
    'du': int
})

额外优化建议

  • 指定文件编码:如果文件使用非utf-8编码(如gbk),在open时明确指定,避免编码错误
  • 过滤无效行:保留空行过滤逻辑,避免处理无意义的空白内容
  • 分块处理(可选):如果文件远大于500MB,可以将数据列表分块,每处理N行就写入临时文件,最后合并,进一步降低内存压力

内容的提问来源于stack exchange,提问作者Dhvani Shah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 10:25:15