正则表达式非匹配项匹配:如何捕获未命中前置规则的内容
解决方法:分类所有Wurk ID并捕获未匹配项
你的需求是把当天日志中Wurk:后的8位字符/数字分成四类:以9、8、7开头的,以及剩下的其他情况。原来的代码只捕获了前三种,漏掉了不符合这三个前缀的项,我们可以通过先统一提取所有符合格式的Wurk ID,再分类判断的方式来解决,这样逻辑更清晰,效率也更高。
修改后的完整代码
import time, glob, re logpath = glob.glob('path\\to\\log*.log')[0] readfile = open(logpath, "r") daysdate = time.strftime("%Y-%m-%d") nine = [] eight = [] seven = [] no_match = [] for line in readfile: # 匹配当天日期行中Wurk:后的8位字母/数字组合 match = re.search(r'{}.*Wurk: ([A-Z0-9]{8})'.format(daysdate), line) if match: wurk_id = match.group(1) # 根据开头字符分类 if wurk_id.startswith('9'): nine.append(wurk_id) elif wurk_id.startswith('8'): eight.append(wurk_id) elif wurk_id.startswith('7'): seven.append(wurk_id) else: no_match.append(wurk_id) # 格式化输出 print("\nNine:\n%s\n" % ",\n".join(nine) + "\nEight:\n%s\n" % ",\n".join(eight) + "\nSeven:\n%s\n" % ",\n".join(seven) + "\nNo matches found:\n%s\n" % ",\n".join(no_match))
关键改进点
- 统一提取所有有效Wurk ID:用一个正则
r'{}.*Wurk: ([A-Z0-9]{8})'.format(daysdate)捕获当天行中所有Wurk:后的8位字母/数字,避免多次正则遍历的冗余。 - 分类判断逻辑:对提取到的每个Wurk ID,通过
startswith()判断前缀,分别放入对应列表;不属于前三种的自动进入no_match。 - 更鲁棒的匹配:正则中的
[A-Z0-9]{8}明确匹配8位字母(大写)+数字的组合,覆盖你示例中的A0885521这类情况。
预期输出
运行修改后的代码后,你会得到:
Nine: 98745061, 98885612 Eight: 88855521 Seven: No matches found: A0885521
备选方案:用正则否定前瞻直接捕获未匹配项
如果你坚持要用单独的正则来抓未匹配项,可以在原来的循环中添加一个正则,用否定前瞻排除9、8、7开头的情况:
# 在原来三个finditer之后添加 for match in re.finditer(r'{}.*Wurk: (?!9|8|7)([A-Z0-9]{8})'.format(daysdate), line): no_match.append(match.group(1))
不过这种方法需要对每行做四次正则匹配,效率不如第一种方法高,推荐优先使用分类判断的方式。
内容的提问来源于stack exchange,提问作者JCJ
相关产品推荐
相关产品推荐

