动态正则提取注册表键:完整路径匹配与转义符优化咨询
动态提取Windows注册表键的正则问题解决
场景与原代码
测试文本:
test_str = "监控对应安全工具(如HKLM:\\SOFTWARE\\Microsoft\\AMSI\\Providers)的服务和启动程序相关Windows注册表项/值的删除操作;监控对应安全工具(如HKLM:\\SOFTWARE\\Policies\\Microsoft\\Windows Defender)的服务和启动程序相关Windows注册表项/值的更改操作。"
原处理代码:
import re iocs_found = { 'windows_path': [ r"HKLM:\SOFTWARE\Microsoft\AMSI\Providers", r"HKLM:\SOFTWARE\Policies\Microsoft\Windows Defender" ] } for path_found in iocs_found['windows_path']: path_found = path_found.replace('\\', '\\\\') print(path_found) regex_pattern = f"[A-Z]+(?:{path_found})" matches = re.findall(regex_pattern, test_str) print(matches) print('\n')
原输出:
HKLM:\\SOFTWARE\\Microsoft\\AMSI\\Providers ['HKLM:\\SOFTWARE\\Microsoft\\AMSI\\Providers'] HKLM:\\SOFTWARE\\Policies\\Microsoft\\Windows Defender ['HKLM:\\SOFTWARE\\Policies\\Microsoft\\Windows']
问题与解决方案
问题1:匹配结果被截断,无法完整获取带空格的注册表路径
原正则中的[A-Z]+仅匹配大写字母序列,遇到Windows Defender里的空格时就停止匹配,导致路径被截断。
解决方法:
修改正则逻辑,用正向预查限定路径的结束边界(空格、右括号、分号这类路径后的分隔符),同时确保路径本身能被完整匹配。结合re.escape()处理路径(见问题2),最终正则可以写成:
regex_pattern = fr"{escaped_path}(?=[\s);])"
(?=[\s);])是正向预查,匹配路径后跟着空格、右括号或分号的位置,保证完整匹配目标路径。
问题2:手动处理双重转义符易出错
原代码中手动替换\为\\来适配正则语法,不仅繁琐还容易遗漏特殊字符。
解决方法:
使用Python内置的re.escape()函数,它会自动转义正则表达式中的所有特殊字符(包括反斜杠、空格等),无需手动处理转义。
优化后完整代码
import re test_str = "监控对应安全工具(如HKLM:\\SOFTWARE\\Microsoft\\AMSI\\Providers)的服务和启动程序相关Windows注册表项/值的删除操作;监控对应安全工具(如HKLM:\\SOFTWARE\\Policies\\Microsoft\\Windows Defender)的服务和启动程序相关Windows注册表项/值的更改操作。" iocs_found = { 'windows_path': [ r"HKLM:\SOFTWARE\Microsoft\AMSI\Providers", r"HKLM:\SOFTWARE\Policies\Microsoft\Windows Defender" ] } for path_found in iocs_found['windows_path']: escaped_path = re.escape(path_found) regex_pattern = fr"{escaped_path}(?=[\s);])" matches = re.findall(regex_pattern, test_str) print(f"匹配结果: {matches}")
优化后输出
匹配结果: ['HKLM:\\SOFTWARE\\Microsoft\\AMSI\\Providers'] 匹配结果: ['HKLM:\\SOFTWARE\\Policies\\Microsoft\\Windows Defender']
内容的提问来源于stack exchange,提问作者Erin Hwang
相关产品推荐
相关产品推荐

