You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用正则表达式提取指定关键词间的完整文本片段

提取number标识间完整文本的正则方案

原有正则的问题

你之前写的(numbers? (\d+|next)[,.]?\s?(\w+))存在两个核心问题:

  • 用\w+仅能匹配连续的单词字符,碰到空格、标点就会停止,自然只能拿到分句的第一个单词
  • 没有设置匹配终止边界,无法在碰到下一个number关键词时自动停止匹配,会出现匹配过短或者过长的问题

可用的正确正则写法

使用非贪婪匹配+正向零宽断言设置边界,正则模式如下:

r'number\s+(\d+|next)[,.]?\s*(.*?)(?=\s*number|$)'

各部分作用说明:

  • number\s+:定位number关键词,匹配关键词后跟随的空白字符
  • (\d+|next):捕获第一个分组,即number后紧跟的数字/next标识
  • [,.]?\s*:匹配标识后可能出现的逗号、句号、空白字符
  • (.*?):非贪婪模式匹配任意长度字符,即两个number之间的分句内容
  • (?=\s*number|$):正向零宽断言作为匹配终止边界:当匹配到下一个前置任意空白的number关键词,或者到达字符串末尾时就停止匹配,不会把下一段标识的内容纳入当前匹配结果

代码运行示例

import re

source_text = "The conditions are:  number 1, the patient is allergic to dust, number next, the patient has bronchitis, number 4, The patient heart rate is high."
pattern = r'number\s+(\d+|next)[,.]?\s*(.*?)(?=\s*number|$)'
match_res = re.findall(pattern, source_text)

# 拼接标识和后续内容,去掉末尾多余句号
final_res = [f"{tag}{content.rstrip('.')}" for tag, content in match_res]
print(final_res)

运行输出结果和预期完全一致:

[
    '1, the patient is allergic to dust, ',
    'next, the patient has bronchitis, ',
    '4, The patient heart rate is high'
]

内容的提问来源于stack exchange,提问作者mgilsn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 00:06:24