正则表达式在Regex101可用但Python中返回空列表问题排查
问题排查:Python正则提取配置块返回空列表
问题描述
需要用Python遍历文件列表,提取以pon-onu-mng gpon-onu开头、直至#符号的配置块(包含首尾标识)。配置文件格式示例:
pon-onu-mng gpon-onu_1/2/2:8 ip-host 2 dhcp-enable true ping-response true traceroute-response true dhcp-ip ethuni eth_0/1 from-internet dhcp-ip ethuni eth_0/2 from-internet wifi disable # pon-onu-mng gpon-onu_1/2/2:9 ip-host 2 dhcp-enable true ping-response true traceroute-response true dhcp-ip ethuni eth_0/1 from-internet dhcp-ip ethuni eth_0/2 from-internet wifi disable # pon-onu-mng gpon-onu_1/2/2:10 ip-host 2 dhcp-enable true ping-response true traceroute-response true dhcp-ip ethuni eth_0/1 from-internet dhcp-ip ethuni eth_0/2 from-internet wifi disable #
使用正则表达式pon-onu-mng.gpon-onu(.*?|[\S\s]?)#在测试工具上可正常匹配,但在Python中调用re.findall时返回空列表,已尝试添加re.MULTILINE和re.DOTALL标志,当前代码:
#! /usr/bin/python3 import os from optparse import OptionParser import re import pandas as pd configFileListRaw = os.system('find . -name startrun.dat > configFileList.txt') configFileListRaw = open('configFileList.txt').readlines() os.system('rm configFileList.txt') pattern = re.compile(r'pon-onu-mng.gpon-onu(.*?|[\S\s]?)#') #match = re.findall(pattern,"pon-onu-mng gpon-onu") print(pattern) for i in configFileListRaw: file = open (str(i).strip("\n")) match = re.findall(r'pon-onu-mng.gpon-onu(.*?|[\S\s]?)#',file.read(),re.MULTILINE | re.DOTALL) print (match)
问题原因及解决方案
1. 正则表达式的核心问题
原正则存在两个致命问题:
- 冗余错误的分组逻辑:
(.*?|[\S\s]?)是完全多余的分支匹配——.*?在re.DOTALL模式下已经能匹配包括换行在内的任意字符,而[\S\s]?仅匹配单个字符或空值,这导致分组无法捕获完整的配置内容。 - findall的分组优先级:
re.findall会优先返回正则中捕获分组的内容,而非整个匹配结果。原正则的分组仅尝试捕获开头标识后的片段,且写法错误导致无有效匹配,最终返回空列表。
2. 修正后的正则表达式
使用以下正则可正确捕获完整的配置块(包含首尾标识):
r'(pon-onu-mng gpon-onu.*?#)'
- 用括号包裹整个目标匹配内容,确保
findall返回完整配置块。 - 用
.*?(非贪婪匹配)在re.DOTALL模式下匹配开头到#之间的所有内容,避免多余分支逻辑。
3. 修正后的代码
#! /usr/bin/python3 import os import re # 替代os.system的Pythonic文件查找方式,避免临时文件 configFileList = [] for root, dirs, files in os.walk('.'): for file in files: if file == 'startrun.dat': configFileList.append(os.path.join(root, file)) # 编译正则,仅需re.DOTALL(MULTILINE无必要,未用到行首/行尾匹配) pattern = re.compile(r'(pon-onu-mng gpon-onu.*?#)', re.DOTALL) for file_path in configFileList: with open(file_path, 'r') as f: content = f.read() matches = pattern.findall(content) print(matches)
关键优化说明
- 移除
re.MULTILINE:未使用行首^或行尾$匹配,re.DOTALL已足够让.匹配换行符。 - 用
os.walk替代os.system执行find命令,避免创建临时文件,更安全高效。 - 使用
with语句管理文件打开,自动处理资源释放,避免文件泄漏。
内容的提问来源于stack exchange,提问作者kirkofthefleet
相关产品推荐
相关产品推荐

