You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则捕获所有重复子模式并解决匹配结果的None值问题

解决方案

Python标准库re的重复捕获组只会保留最后一次匹配的结果,第三方regex库的捕获栈返回多条结果是因为重复组匹配了多次,会把每一次匹配单独返回,所以才会出现大量None值。不需要强行用单条正则递归匹配,分两步处理更稳定,也能兼容所有格式的日志,包括你提到的异常样本。


完整实现代码

import re

# 读取摘要内容
with open("你的摘要文件路径.txt", "r", encoding="utf-8") as f:
    summary = f.read()

# 1. 拆分每个文件的独立区块(按空行分割)
file_blocks = [block.strip() for block in summary.split("\n\n") if block.strip()]

# 预编译正则
# 匹配公共字段的正则
common_pattern = re.compile(
    r"File Name: (.+)\n"
    r"File Start Time: (.+)\n"
    r"File End Time: (.+)\n"
    r"Number of Seizures in File: (\d+)"
)
# 匹配单条癫痫发作时间的正则
seizure_pattern = re.compile(
    r"Seizure(?: \d+)? Start Time: (\d+) seconds\n"
    r"Seizure(?: \d+)? End Time: (\d+) seconds"
)

result = []
for block in file_blocks:
    # 先匹配公共字段
    common_match = common_pattern.search(block)
    if not common_match:
        continue
    file_name, start_time, end_time, seizure_cnt = common_match.groups()
    seizure_cnt = int(seizure_cnt)
    # 匹配所有癫痫条目
    seizures = []
    if seizure_cnt > 0:
        for sz_match in seizure_pattern.finditer(block):
            sz_start, sz_end = sz_match.groups()
            seizures.append((int(sz_start), int(sz_end)))
    # 组装结果
    result.append({
        "file_name": file_name,
        "start_time": start_time,
        "end_time": end_time,
        "seizure_count": seizure_cnt,
        "seizures": seizures
    })

# 打印测试
for item in result:
    print(item)

效果说明

  • 无任何None值,每个文件的公共字段和所有癫痫条目都会完整保留
  • 兼容带编号和不带编号的seizure条目,比如单条发作时的Seizure Start Time和多条时的Seizure 1 Start Time都能正常匹配
  • 兼容不同换行格式、空行分隔的区块,适配各类异常格式的日志样本
  • 如果需要输出成和你之前一致的元组格式,只要调整组装逻辑即可,示例如下:
# 输出为每个发作条目都带公共字段的元组格式
output = []
for item in result:
    if not item["seizures"]:
        output.append((item["file_name"], item["start_time"], item["end_time"], item["seizure_count"], None, None))
    else:
        for sz_start, sz_end in item["seizures"]:
            output.append((item["file_name"], item["start_time"], item["end_time"], item["seizure_count"], sz_start, sz_end))

for line in output:
    print(line)

内容的提问来源于stack exchange,提问作者Farhood ET

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 19:54:03