如何用正则捕获所有重复子模式并解决匹配结果的None值问题
解决方案
Python标准库re的重复捕获组只会保留最后一次匹配的结果,第三方regex库的捕获栈返回多条结果是因为重复组匹配了多次,会把每一次匹配单独返回,所以才会出现大量None值。不需要强行用单条正则递归匹配,分两步处理更稳定,也能兼容所有格式的日志,包括你提到的异常样本。
完整实现代码
import re # 读取摘要内容 with open("你的摘要文件路径.txt", "r", encoding="utf-8") as f: summary = f.read() # 1. 拆分每个文件的独立区块(按空行分割) file_blocks = [block.strip() for block in summary.split("\n\n") if block.strip()] # 预编译正则 # 匹配公共字段的正则 common_pattern = re.compile( r"File Name: (.+)\n" r"File Start Time: (.+)\n" r"File End Time: (.+)\n" r"Number of Seizures in File: (\d+)" ) # 匹配单条癫痫发作时间的正则 seizure_pattern = re.compile( r"Seizure(?: \d+)? Start Time: (\d+) seconds\n" r"Seizure(?: \d+)? End Time: (\d+) seconds" ) result = [] for block in file_blocks: # 先匹配公共字段 common_match = common_pattern.search(block) if not common_match: continue file_name, start_time, end_time, seizure_cnt = common_match.groups() seizure_cnt = int(seizure_cnt) # 匹配所有癫痫条目 seizures = [] if seizure_cnt > 0: for sz_match in seizure_pattern.finditer(block): sz_start, sz_end = sz_match.groups() seizures.append((int(sz_start), int(sz_end))) # 组装结果 result.append({ "file_name": file_name, "start_time": start_time, "end_time": end_time, "seizure_count": seizure_cnt, "seizures": seizures }) # 打印测试 for item in result: print(item)
效果说明
- 无任何None值,每个文件的公共字段和所有癫痫条目都会完整保留
- 兼容带编号和不带编号的seizure条目,比如单条发作时的
Seizure Start Time和多条时的Seizure 1 Start Time都能正常匹配 - 兼容不同换行格式、空行分隔的区块,适配各类异常格式的日志样本
- 如果需要输出成和你之前一致的元组格式,只要调整组装逻辑即可,示例如下:
# 输出为每个发作条目都带公共字段的元组格式 output = [] for item in result: if not item["seizures"]: output.append((item["file_name"], item["start_time"], item["end_time"], item["seizure_count"], None, None)) else: for sz_start, sz_end in item["seizures"]: output.append((item["file_name"], item["start_time"], item["end_time"], item["seizure_count"], sz_start, sz_end)) for line in output: print(line)
内容的提问来源于stack exchange,提问作者Farhood ET
相关产品推荐
相关产品推荐

