如何提取字符串中包含数字的句子片段并定位对应索引
提取含数字的文本片段的实现方法
核心思路
你已经掌握数字提取方法的前提下,只需要额外拿到数字在原字符串中的位置索引,再根据你需要的片段边界(分句/固定长度上下文)计算目标片段的起止索引,直接对原字符串切片即可。re.finditer 会返回所有匹配结果的内容和对应的在原字符串中的起止位置,你可以直接通过返回的span()方法获取索引,不需要自己手动计算位置。
方案1:提取数字所在的完整分句(匹配你给出的示例场景)
你的示例中是按逗号拆分两个分句,我们可以用正则匹配所有分句,同时记录每个分句的位置,筛选出包含数字的分句即可:
import re raw_text = "This is a sentence with numbers, and this is not a sentence with numbers because 123." # 匹配分句,这里分隔符设为中文/英文的逗号、句号、感叹号、问号,可按需调整 clause_pattern = re.compile(r'[^,。!?,.!?]+') for match in clause_pattern.finditer(raw_text): clause = match.group().strip() # 你已有判断是否含数字的逻辑,这里直接复用即可,示例用正则简单判断 if re.search(r'\d', clause): print(clause) # 输出结果:and this is not a sentence with numbers because 123 # 如需获取索引直接调用match.span(),示例中返回结果为(33, 94),用raw_text[33:94]即可切片提取
方案2:提取数字前后固定长度的上下文片段
如果不需要按分句拆分,只需要拿数字前后N个字符的片段,直接用数字的索引计算即可:
import re raw_text = "This is a sentence with numbers, and this is not a sentence with numbers because 123." # 匹配数字,拿到数字的起止索引 num_match = re.search(r'\d+', raw_text) if num_match: num_start, num_end = num_match.span() # 前后各保留50个字符,可按需调整,用max/min避免索引越界 context_len = 50 fragment_start = max(0, num_start - context_len) fragment_end = min(len(raw_text), num_end + context_len) target_fragment = raw_text[fragment_start:fragment_end] print(target_fragment) # 输出结果:, and this is not a sentence with numbers because 123.
如果需要处理文本中存在多个数字的场景,把re.search替换为re.finditer遍历所有匹配结果即可。
内容的提问来源于stack exchange,提问作者Tab1e
相关产品推荐
相关产品推荐

