如何使用Python提取段落中包含0-85之间数值的句子,含正则实现方案
解决方案
下面提供两种可直接在Python中运行的实现思路,均可满足提取含0~85区间数值句子的需求:
方案1:纯正则表达式实现
首先构造匹配0~85区间整数的正则规则:
- 匹配逻辑:覆盖1位数字(09)、十位为07的两位数(0079)、十位为8且个位为05的两位数(80~85),加入整词边界避免匹配长数字中的片段。
- 范围匹配正则片段:
\b(?:[0-7]?\d|8[0-5])\b
完整实现代码:
import re text = """The patient is suffering from fever. Their relatives come to visit them. The patient age is 20 year. His brother could not visit him due to some other work. There is another patient whose age is 30 year old. The second patient is watching him from window.""" # 完整匹配规则:提取包含符合范围数值的句子,自动去除句尾的句号和前后空格 pattern = re.compile(r'([^.]*(?:\b(?:[0-7]?\d|8[0-5])\b)[^.]*)\.') result = [s.strip() for s in pattern.findall(text)] print(result)
运行输出和预期结果完全一致:
['The patient age is 20 year', 'There is another patient whose age is 30 year old']
方案2:分句+数值校验的混合实现(更推荐)
纯正则的范围匹配灵活性较差,若后续需要调整数值阈值需要重写规则,更推荐先拆分所有句子,再逐句校验数值范围的实现方式:
import re text = """The patient is suffering from fever. Their relatives come to visit them. The patient age is 20 year. His brother could not visit him due to some other work. There is another patient whose age is 30 year old. The second patient is watching him from window.""" # 先按句号拆分所有句子,过滤空串并去除前后空格 sentences = [s.strip() for s in text.split('.') if s.strip()] result = [] for sent in sentences: # 提取句子中的所有整数 nums = re.findall(r'\b\d+\b', sent) # 校验是否存在落在0~85区间的数值 for num_str in nums: if 0 <= int(num_str) <= 85: result.append(sent) break print(result)
该方案的优势是逻辑可读性更高,调整数值范围仅需要修改判断条件即可,适配性更强。
内容的提问来源于stack exchange,提问作者XGB
相关产品推荐
相关产品推荐

