使用Python格式化解析日志文件并转换为CSV
日志提取并生成CSV的解决方案
问题背景
我最初尝试用PowerShell解决日志提取并生成CSV的问题但未成功,转而使用Python。目前已能将指定表头写入CSV文件,但不清楚如何从目标日志文件中提取所需的特定数据项(无需日志自带表头)。
现有代码
import csv import numpy as np header = np.asarray([["Connection","IKE PEER","TYPE","REKEY","ENCRYPT","AUTH","ROLE","STATE","HASH","LIFETIME","LIFETIME REMAINING"]]) with open('output.csv', 'w') as f: mywriter = csv.writer(f, delimiter=',') mywriter.writerows(header)
待解析日志示例
1 IKE Peer: 184.188.106.30 (PHX2) Type : L2L Role : initiator Rekey : no State : MM_ACTIVE Encrypt : 3des Hash : SHA Auth : preshared Lifetime: 86400 Lifetime Remaining: 52140 2 IKE Peer: 96.39.70.182 (Worcester) Type : L2L Role : initiator Rekey : no State : MM_ACTIVE Encrypt : aes-256 Hash : SHA Auth : preshared Lifetime: 28800 Lifetime Remaining: 5929 ...(其余日志内容略)
期望生成的CSV格式
生成包含对应数据项的表格,每行对应一个连接的完整数据,列顺序与定义的表头一致:
| Connection | IKE PEER | TYPE | REKEY | ENCRYPT | AUTH | ROLE | STATE | HASH | LIFETIME | LIFETIME REMAINING |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 184.188.106.30 (PHX2) | L2L | no | 3des | preshared | initiator | MM_ACTIVE | SHA | 86400 | 52140 |
| 2 | 96.39.70.182 (Worcester) | L2L | no | aes-256 | preshared | initiator | MM_ACTIVE | SHA | 28800 | 5929 |
实现代码
可以通过逐行读取日志、按连接块分组提取字段的方式完成需求,无需依赖numpy,直接用csv模块更简洁:
import csv # 定义表头 header = ["Connection","IKE PEER","TYPE","REKEY","ENCRYPT","AUTH","ROLE","STATE","HASH","LIFETIME","LIFETIME REMAINING"] # 存储所有连接数据 connections = [] # 读取日志文件(替换为你的日志路径) with open('log.txt', 'r') as log_file: current_conn = {} for line in log_file: line = line.strip() if not line: continue # 识别连接起始行(以数字开头) if line[0].isdigit(): # 保存上一个连接的数据 if current_conn: connections.append(current_conn) current_conn = {} # 拆分连接编号和IKE Peer信息 num_part, peer_part = line.split('IKE Peer:', 1) current_conn['Connection'] = num_part.strip() current_conn['IKE PEER'] = peer_part.strip() else: # 处理一行多字段的情况(如Type和Role在同一行) field_groups = [group.strip() for group in line.split(' ') if group.strip()] for group in field_groups: key, value = group.split(':', 1) key = key.strip().upper() # 匹配表头的字段名 if key == 'LIFETIME REMAINING': current_conn[key] = value.strip() else: current_conn[key] = value.strip() # 保存最后一个连接的数据 if current_conn: connections.append(current_conn) # 写入CSV文件 with open('output.csv', 'w', newline='') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=header) writer.writeheader() writer.writerows(connections)
代码说明
- 按数字起始行划分每个连接的数据块,避免跨块混淆
- 自动处理一行包含多个字段的情况,拆分后分别提取键值对
- 统一字段名格式,确保与表头完全匹配
- 使用
csv.DictWriter自动按表头顺序写入数据,无需手动对应列位置
内容的提问来源于stack exchange,提问作者Jamie Lovenduski
相关产品推荐
相关产品推荐

