You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式提取指定协议组合的Wireshark日志关键内容

针对Wireshark文本日志的XML协议帧提取正则方案

核心正则表达式(Python兼容)

以下正则会精准匹配协议链为eth:ethertype:ip:tcp:http:xml的帧,并提取你需要的4项内容:

import re

pattern = re.compile(
    r'^Frame \d+: .*?eth:ethertype:ip:tcp:http:xml$\n.*?(Internet Protocol Version \d, src: .*? \(.*?\), dst: .*? \(.*?\))\n.*?(Transmission Control Protocol, src port: .*?, dst port: .*?)\n.*?Hypertext Transfer Protocol\s*\n\s*(.*?)\s*\n.*?(eXtensible Markup Language.*?(?:</SOAP-ENV:Envelope>|</s:Envelope>))',
    re.DOTALL | re.MULTILINE
)

代码使用示例

# 读取Wireshark导出的纯文本日志
with open('wireshark_export.txt', 'r', encoding='utf-8') as f:
    log_data = f.read()

# 提取所有符合条件的帧数据
results = pattern.findall(log_data)

# 遍历输出结果
for i, (ip_line, tcp_line, http_line, xml_content) in enumerate(results, 1):
    print(f"--- 第{i}个匹配帧 ---")
    print(f"IP信息: {ip_line}")
    print(f"TCP端口: {tcp_line}")
    print(f"HTTP行: {http_line.strip()}")
    print(f"XML内容:\n{xml_content}\n")

正则逻辑说明(解决之前匹配失败的关键)

  1. 帧过滤:^Frame \d+: .*?eth:ethertype:ip:tcp:http:xml$ 精准定位目标协议链的帧开头,避免匹配其他无关帧。
  2. IP行捕获:(Internet Protocol Version \d, src: .*? \(.*?\), dst: .*? \(.*?\)) 完整捕获包含源/目的IP的整行内容,非贪婪匹配避免跨内容。
  3. TCP行捕获:(Transmission Control Protocol, src port: .*?, dst port: .*?) 捕获包含源/目的端口的TCP行。
  4. HTTP行捕获:跳过Hypertext Transfer Protocol行,直接提取下一行的实际HTTP请求/响应状态行(处理了换行和空白字符)。
  5. XML内容捕获:开启re.DOTALL后,.*?会跨行匹配直到遇到指定的SOAP闭合标签,兼容</SOAP-ENV:Envelope>和</s:Envelope>两种常见格式。

常见问题排查(之前匹配失败的可能原因)

  • 未开启re.DOTALL:导致正则无法跨行匹配XML或帧内的多行内容。
  • 未精准过滤协议链:匹配了非目标协议的帧,导致提取内容混乱。
  • XML闭合标签未兼容多种格式:只写了一种SOAP信封闭合标签,漏掉了另一种。
  • 未处理空白字符:HTTP行前后的换行、空格导致匹配失败。

内容的提问来源于stack exchange,提问作者RKIDEV

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 20:38:10