如何正确使用Regex的Match或Group方法解析日志并导出为CSV
日志错误记录解析导出CSV实现方法
需求说明
从给定格式的日志文件中筛选错误告警记录,解析后导出包含Date、Time、ErrorLevel、InstrumentType、Value、ContactInfo、Description共7个字段的CSV文件。
日志样例
[2019-Apr-01 13:00:02.343][Test][Info][2929:12][Writing.To.Log] Reading ABC. [2019-Apr-01 13:00:03.343][Test][Alarm][8192:12][Test.In.Progress] Severity: Error0, Description: 'Information = ABC:1:ABC Value went over upper limit, Value = 93.0625 C, Error0UCL = 30 C.', ContactInfo: 555-555-5555 [2019-Apr-01 13:00:04.343][Test][Info][2929:12][Writing.To.Log] Reading DEF. [2019-Apr-01 13:00:05.353][Test][Alarm][8193:12][Test.In.Progress] Severity: Error0, Description: 'Information = DEF:1:DEF Value went over upper limit, Value = 93.0625 C, Error0UCL = 30 C.', ContactInfo: 555-555-5555
期望输出格式
| Date | Time | ErrorLevel | InstrumentType | Value | ContactInfo | Description |
|---|---|---|---|---|---|---|
| 2019-Apr-01 | 13:00:03.343 | 0 | ABC | 93.0625 | 555-555-5555 | Information = ABC:1:ABC Value went over upper limit, Value = 93.0625 C, Error0UCL = 30 C. |
| 2019-Apr-01 | 13:00:05.353 | 0 | DEF | 93.0625 | 555-555-5555 | Information = DEF:1:DEF Value went over upper limit, Value = 93.0625 C, Error0UCL = 30 C. |
完整实现代码
import re import csv # 主正则:一次性匹配日期、时间、错误级别、描述内容、联系方式 main_pattern = re.compile(r'^\[(\d{4}-[A-Za-z]{3}-\d{2}) (\d{2}:\d{2}:\d{2}\.\d{3})\].*?Severity: Error(\d+), Description: \'([^\']+)\', ContactInfo: (\d{3}-\d{3}-\d{4})') # 从描述中提取设备类型和数值的正则 instrument_pattern = re.compile(r'Information = (.*?):') value_pattern = re.compile(r'Value = ([\d.]+)') with open('Parsed.csv', 'w', newline='', encoding='utf-8') as out_file: writer = csv.writer(out_file) writer.writerow(['Date', 'Time', 'ErrorLevel', 'InstrumentType', 'Value', 'ContactInfo', 'Description']) with open('Log.txt', 'r', encoding='utf-8') as in_file: for line in in_file: line = line.strip() main_match = main_pattern.match(line) if main_match: date = main_match.group(1) time = main_match.group(2) error_level = main_match.group(3) description = main_match.group(4) contact = main_match.group(5) # 提取设备类型 ins_match = instrument_pattern.search(description) instrument = ins_match.group(1) if ins_match else '' # 提取数值 val_match = value_pattern.search(description) value = val_match.group(1) if val_match else '' # 写入csv writer.writerow([date, time, error_level, instrument, value, contact, description])
实现逻辑说明
- 主正则利用分组匹配一次性提取日志行的5个核心字段,避免多次正则匹配的性能损耗
- 新增
newline=''参数避免CSV写入时出现多余空行,指定编码兼容多语言环境 - 对从描述中提取的字段做兼容处理,匹配失败时写入空值避免程序异常
内容的提问来源于stack exchange,提问作者QueenBee
相关产品推荐
相关产品推荐

