请求协助编写Python脚本实现日志文件转CSV
日志转CSV的Python脚本问题
我需要编写一个Python脚本,将指定日志文件转换为CSV格式。以下是输入的日志内容:
Exception: Cannot open file "C:\ProgramData\Presence\Log\pco_nhp01_CT_1800.log". Access is denied Original message: [27/01/2023 12:37:44:675] TID:[14588] ENTER FUNCTION SetActive :: = { Value: True, } [27/01/2023 12:37:44:675] TID:[14588] VERBOSE GetServerIP : 10.10.10.155 [27/01/2023 12:37:44:691] TID:[14588] LEAVE FUNCTION SetActive :: = { Active: True, } [27/01/2023 12:37:44:694] TID:[14588] ENTER FUNCTION SetActive :: = { Value: True, } [27/01/2023 12:37:44:694] TID:[14588] VERBOSE GetServerIP : 10.10.10.155 [27/01/2023 12:37:44:703] TID:[14588] LEAVE FUNCTION SetActive :: = { Active: True, } [27/01/2023 12:37:44:703] TID:[14588] ENTER FUNCTION MonitorDevice :: = { Device: 201122, } [27/01/2023 12:37:44:707] TID:[7060] ENTER FUNCTION TEventsManager.AddEvent :: = { ACSTAEvent: CSTACONFIRMATION CSTAR_MONITORS_CON, CTIRequestID: 2, } [27/01/2023 12:37:53:711] TID:[7060] LEAVE FUNCTION TEventsManager.AddEvent
我尝试了以下代码,但无法正确提取日志信息,达不到预期输出:
import csv with open('pco_nhp01_CT_1800.log', 'r') as log_file: log_data = log_file.readlines() with open('logfile.csv', 'w', newline='') as csv_file: writer = csv.writer(csv_file) writer.writerow(['Datetime', 'TID', 'Message']) for line in log_data: if line.startswith('['): parts = line.split(']') datetime = parts[0][1:] tid = parts[1][6:] message = parts[2][1:] writer.writerow([datetime, tid, message])
预期输出说明
预期的CSV文件包含三列:Datetime、TID、Message。其中:
- 每条日志的时间和TID对应开头带
[的行内容 Message字段需要包含该日志条目下的所有后续内容,直到遇到空行或下一条日志的时间行。比如第一条日志的Message要包含ENTER FUNCTION、SetActive :: =以及大括号内的内容;VERBOSE条目要包含VERBOSE和GetServerIP : 10.10.10.155等。同时开头的异常信息也需要作为单独条目处理。
修正后的代码
import csv import re def parse_log_to_csv(log_path, csv_path): # 正则匹配日志头行:[时间] TID:[线程ID] log_header_pattern = re.compile(r'\[([^\]]+)\] TID:\[(\d+)\]') current_datetime = None current_tid = None current_message = [] with open(log_path, 'r', encoding='utf-8') as log_file: lines = [line.strip() for line in log_file if line.strip()] with open(csv_path, 'w', newline='', encoding='utf-8') as csv_file: writer = csv.writer(csv_file) writer.writerow(['Datetime', 'TID', 'Message']) # 先处理开头的异常行 if lines and not log_header_pattern.match(lines[0]): # 检查是否有Original message行 if len(lines) > 1 and lines[1].startswith('Original message:'): match = log_header_pattern.search(lines[1]) if match: current_datetime = match.group(1) current_tid = match.group(2) current_message.append(lines[0]) # 跳过Original message行,从下一行开始收集内容 start_idx = 2 else: # 无匹配的话,异常行单独作为条目 writer.writerow(['', '', lines[0]]) start_idx = 1 else: writer.writerow(['', '', lines[0]]) start_idx = 1 else: start_idx = 0 for line in lines[start_idx:]: header_match = log_header_pattern.match(line) if header_match: # 有之前的未写入内容,先写入 if current_datetime and current_tid and current_message: full_message = '\n'.join(current_message) writer.writerow([current_datetime, current_tid, full_message]) # 更新当前头信息 current_datetime = header_match.group(1) current_tid = header_match.group(2) current_message = [] else: # 收集消息内容 current_message.append(line) # 写入最后一条日志 if current_datetime and current_tid and current_message: full_message = '\n'.join(current_message) writer.writerow([current_datetime, current_tid, full_message]) # 调用函数 parse_log_to_csv('pco_nhp01_CT_1800.log', 'logfile.csv')
代码说明
- 使用正则表达式匹配日志的时间和TID行,比字符串分割更稳定
- 处理开头的异常信息,确保不会丢失这部分内容
- 逐行收集每条日志对应的多行消息内容,直到遇到下一条日志的头行或文件结束
- 将收集到的多行消息合并为一个字符串,写入CSV的Message列
- 过滤空行,只保留有效内容行
内容的提问来源于stack exchange,提问作者N K
相关产品推荐
相关产品推荐

