You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何将无固定分隔符的日志文本解析为pandas DataFrame?

日志结构化方案

直接用正则命名捕获组就可以实现需求,不需要你从头学正则,直接用下面的现成代码即可:

核心思路

我们用带命名捕获组的正则表达式匹配每行日志的对应字段,匹配不到的可选字段(比如START/END行没有的时长、内存指标)会自动返回空值,刚好适配不同行结构不同的场景。

完整实现代码

import re
import pandas as pd
import os

# 预编译正则表达式,适配你给出的日志格式
log_pattern = re.compile(
    r'^(?P<DateTime>\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{3}Z) '
    r'(?P<Keyword>START|END|REPORT) '
    r'RequestId: (?P<RequestId>\w+)'
    r'(?:.*Duration: (?P<Duration>\d+\.?\d*) ms)?'
    r'(?:.*Billed Duration: (?P<BilledDuration>\d+) ms)?'
    r'(?:.*Memory Size: (?P<MemorySize>\d+) MB)?'
    r'(?:.*Max Memory Used: (?P<MaxMemoryUsed>\d+) MB)?'
)

log_folder = "替换为你的txt日志所在的文件夹路径"
parsed_result = []

# 批量读取所有日志文件
for file_name in os.listdir(log_folder):
    if not file_name.endswith(".txt"):
        continue
    with open(os.path.join(log_folder, file_name), 'r', encoding='utf-8') as f:
        for line in f:
            stripped_line = line.strip()
            if not stripped_line:
                continue
            match_res = log_pattern.match(stripped_line)
            if match_res:
                parsed_result.append(match_res.groupdict())

# 转换为DataFrame并做类型转换
log_df = pd.DataFrame(parsed_result)
log_df['DateTime'] = pd.to_datetime(log_df['DateTime'])
num_cols = ['Duration', 'BilledDuration', 'MemorySize', 'MaxMemoryUsed']
log_df[num_cols] = log_df[num_cols].apply(pd.to_numeric, errors='coerce')

后续处理说明

你要做内存异常检测时,直接过滤REPORT类型的日志即可拿到所有带内存指标的记录:

report_logs = log_df[log_df['Keyword'] == 'REPORT'].reset_index(drop=True)

拿到的report_logs可以直接用于3σ阈值检测、孤立森林异常检测等场景。

内容的提问来源于stack exchange,提问作者Kami

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 22:48:04