用Python将log.txt转换为JSON时遇解析问题求助
问题背景
我正在学习Python,想把系统日志log.txt转换为JSON格式,要求每个事件作为独立对象,将LogName、MachineName等字段拆分为键值对,但当前脚本输出不符合预期。
当前使用的Python脚本
import re import json import os def parse_log_file(input_file): events = [] with open(input_file, 'r') as file: log_content = file.read() # extract individual events event_pattern = re.compile(r'Event \d+\s+(.*?)\s+(?=(?:Event \d+|$))', re.DOTALL) matches = event_pattern.findall(log_content) for match in matches: event_dict = {} lines = match.split('\n') for line in lines: if line.strip(): key, value = map(str.strip, line.split(':', 1)) event_dict[key] = value events.append(event_dict) # Write the JSON output with the same name as the input file output_file = os.path.splitext(input_file)[0] + ".json" with open(output_file, 'w') as json_file: json.dump(events, json_file, indent=4) print(f"JSON file saved as,{output_file}") if __name__ == "__main__": input_file = "log.txt" parse_log_file(input_file)
期望输出格式
Event 1
{ "LogName" : "System",
"MachineName" : "LAPTOP" ,
"ProviderName" : "Intel",
"LevelDisplayName" : "Information",
"Message" : "Check the remaining resource budget. Module exceeds resource budget, failed to AllocateFwCps,
STATUS = Insufficient system resources exist to complete the API.." },Event 2 {
"LogName" : "System",
"MachineName" : "LAPTOP"
"ProviderName" : "Microsoft-Windows-Kernel-Power"
"LevelDisplayName" : "Information"
"Message" : "The system session has transitioned from 186 to 188. Reason InputPoUserPresent
BootId: 67"
}
实际输出情况
LogName : "System
MachineName : LAPTOP
ProviderName : Microsoft-Windows-Kernel-Power
LevelDisplayName : Information
Message : The system session has transitioned from 186 to 188. Reason InputPoUserPresent
BootId: 67"
问题分析
- 多行字段拆分错误:Message等字段包含换行和冒号时,脚本会将后续行误识别为独立键值对,导致字段内容不完整且生成无效键。
- 未捕获事件编号:原正则未提取Event编号,生成的JSON中缺少事件标识信息。
- 正则匹配边界不准确:原正则的空白匹配逻辑可能导致事件内容分割错误,出现字段内容混叠。
修复后的脚本
import re import json import os def parse_log_file(input_file): events = [] with open(input_file, 'r') as file: log_content = file.read() # 修正正则:捕获Event编号和对应内容,准确分割事件边界 event_pattern = re.compile(r'(Event \d+)\s+(.*?)(?=\s*Event \d+|$)', re.DOTALL) matches = event_pattern.findall(log_content) for event_id, event_content in matches: event_dict = {} # 添加事件编号到字典 event_dict["EventId"] = event_id lines = event_content.split('\n') last_key = None for line in lines: stripped_line = line.strip() if not stripped_line: continue # 判断是否为键值对行 if ':' in stripped_line: key_part, value_part = stripped_line.split(':', 1) key = key_part.strip() value = value_part.strip() # 处理多行带引号的字段(如Message) if value.startswith('"') and not value.endswith('"'): last_key = key event_dict[key] = value else: event_dict[key] = value last_key = None else: # 将后续行合并到上一个字段的内容中 if last_key is not None: event_dict[last_key] += '\n' + stripped_line events.append(event_dict) # 生成JSON输出文件 output_file = os.path.splitext(input_file)[0] + ".json" with open(output_file, 'w') as json_file: json.dump(events, json_file, indent=4) print(f"JSON file saved as {output_file}") if __name__ == "__main__": input_file = "log.txt" parse_log_file(input_file)
修复说明
- 正则优化:现在可以准确捕获每个事件的编号和内容,避免事件边界混淆。
- 多行字段处理:识别带引号的多行字段,将后续行合并到同一个键下,保证字段内容完整。
- 添加事件标识:将Event编号存入字典,满足输出中事件标识的需求。
内容的提问来源于stack exchange,提问作者tthhhww

