You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从日志文件提取文本与JSON并关联,生成Pandas DataFrame?

日志提取与转换为Pandas DataFrame方案

实现思路

  • 逐行读取日志,以run作为片段起始标记,捕获该行的时间戳、run编号、用户名
  • 持续读取后续行,直到遇到]结束标记,拼接完整JSON内容
  • 解析JSON数据,将元数据(时间戳、run、user)与JSON字段合并
  • 整理所有数据,转换为指定结构的Pandas DataFrame

代码实现

import pandas as pd
import json
import re

# 替换为你的日志文件路径
log_file = "target.log"

# 存储最终数据的列表
result_data = []

# 正则匹配起始行,提取时间戳、run、user信息
start_line_pattern = re.compile(r"^(\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}), run: (\d+), user: (\w+) json:$")

with open(log_file, 'r', encoding='utf-8') as f:
    current_meta = None
    json_content = []
    for line in f:
        stripped_line = line.strip()
        # 匹配起始行(含run关键字)
        match_res = start_line_pattern.match(stripped_line)
        if match_res:
            # 先处理上一个未完成的片段
            if current_meta and json_content:
                json_str = ''.join(json_content)
                try:
                    json_obj = json.loads(json_str)[0]
                    result_data.append({
                        'timestamp': current_meta[0],
                        'run': int(current_meta[1]),
                        'user': current_meta[2],
                        'value': int(json_obj['value']),
                        'error': int(json_obj['error'])
                    })
                except json.JSONDecodeError:
                    # 忽略解析失败的片段,可按需调整处理逻辑
                    pass
            # 更新当前元数据,重置JSON内容列表
            current_meta = match_res.groups()
            json_content = []
        elif current_meta is not None:
            # 收集JSON行,直到遇到结束符]
            json_content.append(stripped_line)
            if ']' in stripped_line:
                json_str = ''.join(json_content)
                try:
                    json_obj = json.loads(json_str)[0]
                    result_data.append({
                        'timestamp': current_meta[0],
                        'run': int(current_meta[1]),
                        'user': current_meta[2],
                        'value': int(json_obj['value']),
                        'error': int(json_obj['error'])
                    })
                except json.JSONDecodeError:
                    pass
                # 重置状态,准备下一个片段
                current_meta = None
                json_content = []

# 转换为DataFrame并调整列顺序
df = pd.DataFrame(result_data)[['timestamp', 'run', 'user', 'value', 'error']]
print(df)

代码说明

  • 正则匹配:通过start_line_pattern精准定位包含run的起始行,一次性提取所需元数据
  • 片段处理:识别起始行后,持续收集后续行直到]出现,确保获取完整的JSON内容
  • 数据合并:解析JSON后将字段与元数据合并为字典,统一存入列表
  • DataFrame转换:将列表转为DataFrame,并调整列顺序匹配目标结构

测试输出

用你提供的日志示例运行代码,将得到如下结果:

timestamp  run   user  value  error
0  2022-12-15 12:45:06    1  james     30      8
1  2022-12-15 12:47:36    2  kelly     15      3

内容的提问来源于stack exchange,提问作者DJC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 07:10:32