使用Python正则表达式从日志提取网络信息遇阻求助
日志字段提取:IPv4/变量混合场景的解决方案
问题描述
我正尝试从日志文件中提取多个字段,但处理IPv4地址、子网与变量混合的场景时遇到了困难,目前只能匹配单一类型的字段(比如单独的IP或字符串)。
现有代码如下:
import re regex = re.search( r'.*(?P<destination_address>\b((?:\d+\.){3}\d+(?:/\d+)?)|\w+)\b(?P<destination_port>\d+)?\b(?P<destination_options>\w)?(?=via|\Z|//)', "Myfirewall add 50750 set Mycounter allow udp from any to 123.45.67.89/28 123 via someotheriface" ) regex2 = re.search( r'.*(?P<destination_address>\b((?:\d+\.){3}\d+(?:/\d+)?)|\w+)\b(?P<destination_port>\d+)?\b(?P<destination_options>\w)?(?=via|\Z|//)', "Myfirewall add 50750 set Mycounter allow udp from 123.45.67.89/28 to Mynic opt1 opt2,opt3 via someotheriface" )
当前两个正则均无匹配结果,期望达成:
regex.group("destination_port") == "123"regex2.group("destination_options") == "opt1 opt2,opt3"
已能提取“to”之前的所有字段,剩余难点:
- 捕获“to”与“via”、注释起始符(//)或换行符之间的字符串
- 判断内容是常量(IPv4地址/子网)还是变量(字符串)——核心需求
- 分离主内容与次要内容(端口或选项)
若正则实现难度大,接受替代方案。
解决方案
思路解析
原正则的问题在于:
- 贪婪匹配
.*会吃掉“to”前面的内容,导致后续捕获组无法定位目标区域 - 对“选项”的匹配仅限制为单个单词
\w,无法匹配包含空格、逗号的多选项内容 - 未明确限定“to”作为起始锚点,导致匹配范围混乱
我们可以分两步处理:先定位“to”到终止符之间的目标片段,再对片段进行字段提取。
代码实现
方案1:正则表达式优化
import re # 第一步:匹配"to"到终止符之间的内容 target_pattern = re.compile(r'to\s+(.*?)(?=\s+via|\s+//|\n|\Z)', re.DOTALL) # 第二步:分别匹配IP/变量、端口、选项的模式 ip_port_pattern = re.compile(r'(?P<destination_address>\b(?:\d+\.){3}\d+(?:/\d+)?\b)\s+(?P<destination_port>\d+)') var_options_pattern = re.compile(r'(?P<destination_address>\b\w+\b)\s+(?P<destination_options>.*)') # 测试第一个案例 line1 = "Myfirewall add 50750 set Mycounter allow udp from any to 123.45.67.89/28 123 via someotheriface" match_target = target_pattern.search(line1) if match_target: target_content = match_target.group(1).strip() # 先尝试匹配IP+端口 ip_port_match = ip_port_pattern.fullmatch(target_content) if ip_port_match: print("案例1结果:") print(f"destination_address: {ip_port_match.group('destination_address')}") print(f"destination_port: {ip_port_match.group('destination_port')}") # 测试第二个案例 line2 = "Myfirewall add 50750 set Mycounter allow udp from 123.45.67.89/28 to Mynic opt1 opt2,opt3 via someotheriface" match_target2 = target_pattern.search(line2) if match_target2: target_content2 = match_target2.group(1).strip() # 尝试匹配变量+选项 var_options_match2 = var_options_pattern.fullmatch(target_content2) if var_options_match2: print("\n案例2结果:") print(f"destination_address: {var_options_match2.group('destination_address')}") print(f"destination_options: {var_options_match2.group('destination_options')}")
方案2:字符串分割(替代方案)
如果日志格式相对固定,可通过字符串分割简化处理:
import re # 处理第一个案例 line1 = "Myfirewall add 50750 set Mycounter allow udp from any to 123.45.67.89/28 123 via someotheriface" target_part = line1.split('to')[1].split('via')[0].strip().split() if re.fullmatch(r'(?:\d+\.){3}\d+(?:/\d+)?', target_part[0]): print("案例1结果:") print(f"destination_address: {target_part[0]}") print(f"destination_port: {target_part[1]}") # 处理第二个案例 line2 = "Myfirewall add 50750 set Mycounter allow udp from 123.45.67.89/28 to Mynic opt1 opt2,opt3 via someotheriface" target_part2 = line2.split('to')[1].split('via')[0].strip().split(maxsplit=1) if re.fullmatch(r'\w+', target_part2[0]): print("\n案例2结果:") print(f"destination_address: {target_part2[0]}") print(f"destination_options: {target_part2[1]}")
核心需求实现(判断常量/变量)
通过正则直接判断内容类型:
import re def is_ip_cidr(s): return bool(re.fullmatch(r'(?:\d+\.){3}\d+(?:/\d+)?', s)) # 用法示例 addr = "123.45.67.89/28" print(f"{addr} 是常量:{is_ip_cidr(addr)}") addr2 = "Mynic" print(f"{addr2} 是变量:{not is_ip_cidr(addr2)}")
内容的提问来源于stack exchange,提问作者xancho
相关产品推荐
相关产品推荐

