You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3提取Suricata日志特定格式字段的技术求助

处理Suricata日志提取的Python方案

针对你要提取Suricata警报日志里的描述和SID前两段的需求,给你两种简单的实现方法,都是Python 3可用的:

方法一:用字符串分割(适配你尝试的split思路)

这种方法基于日志固定格式拆分字符串,步骤清晰:

# 打开日志文件和输出文件
with open('alerts.log', 'r') as log_file, open('output.txt', 'w') as out_file:
    for line in log_file:
        # 跳过空行
        if not line.strip():
            continue
        # 用[**]分割每行日志,拆分出各核心部分
        parts = line.split('[**]')
        # 处理SID部分:提取[1:2210038:2]里的1:2210038
        sid_raw = parts[1].strip().strip('[]')  # 去掉前后空格和[],得到"1:2210038:2"
        target_sid = ':'.join(sid_raw.split(':')[:2])  # 按:分割后取前两段拼接
        # 处理警报描述:提取中间的SURICATA开头内容
        alert_desc = parts[2].strip()
        # 按要求格式写入输出文件
        out_file.write(f'# {alert_desc}\n')
        out_file.write(f'{target_sid}\n\n')  # 空行分隔每条记录,可根据需求调整

代码说明:

  • line.split('[**]')把日志拆分为多段,我们只需要第1、2段内容
  • sid_raw.split(':')[:2]将完整SID拆分为列表后取前两个元素,再用:拼接,得到目标格式的SID
  • with open会自动处理文件关闭,不用手动写close语句

方法二:用正则表达式(更稳定,适配日志格式小变动)

如果日志里的分隔空格偶尔有变化,正则能更精准匹配目标内容:

import re

# 匹配SID和警报描述的正则规则
alert_pattern = re.compile(r'\[\*\*\] \[(\d+:\d+):\d+\] (.*?) \[\*\*\]')

with open('alerts.log', 'r') as log_file, open('output.txt', 'w') as out_file:
    for line in log_file:
        match_result = alert_pattern.search(line)
        if match_result:
            target_sid = match_result.group(1)  # 提取第一个分组的内容:1:2210038
            alert_desc = match_result.group(2)  # 提取第二个分组的内容:SURICATA STREAM FIN out of window
            out_file.write(f'# {alert_desc}\n')
            out_file.write(f'{target_sid}\n\n')

正则说明:

  • \[\*\*\]匹配[**](因为[]和*是正则特殊字符,需要加反斜杠转义)
  • (\d+:\d+):\d+匹配完整SID,把前两段数字加冒号作为第一个可提取分组
  • (.*?)非贪婪匹配警报描述,避免把后面的分类、优先级内容也包含进来

你可以根据自己的日志格式选其中一种方法,运行后就能得到符合要求的输出内容。

内容的提问来源于stack exchange,提问作者tryin2code

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 03:45:33