You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取字符串中包含数字的句子片段并定位对应索引

提取含数字的文本片段的实现方法

核心思路

你已经掌握数字提取方法的前提下,只需要额外拿到数字在原字符串中的位置索引,再根据你需要的片段边界(分句/固定长度上下文)计算目标片段的起止索引,直接对原字符串切片即可。re.finditer 会返回所有匹配结果的内容和对应的在原字符串中的起止位置,你可以直接通过返回的span()方法获取索引,不需要自己手动计算位置。

方案1:提取数字所在的完整分句(匹配你给出的示例场景)

你的示例中是按逗号拆分两个分句,我们可以用正则匹配所有分句,同时记录每个分句的位置,筛选出包含数字的分句即可:

import re

raw_text = "This is a sentence with numbers, and this is not a sentence with numbers because 123."
# 匹配分句,这里分隔符设为中文/英文的逗号、句号、感叹号、问号,可按需调整
clause_pattern = re.compile(r'[^,。!?,.!?]+')
for match in clause_pattern.finditer(raw_text):
    clause = match.group().strip()
    # 你已有判断是否含数字的逻辑,这里直接复用即可,示例用正则简单判断
    if re.search(r'\d', clause):
        print(clause)
        # 输出结果:and this is not a sentence with numbers because 123
        # 如需获取索引直接调用match.span(),示例中返回结果为(33, 94),用raw_text[33:94]即可切片提取

方案2:提取数字前后固定长度的上下文片段

如果不需要按分句拆分,只需要拿数字前后N个字符的片段,直接用数字的索引计算即可:

import re

raw_text = "This is a sentence with numbers, and this is not a sentence with numbers because 123."
# 匹配数字,拿到数字的起止索引
num_match = re.search(r'\d+', raw_text)
if num_match:
    num_start, num_end = num_match.span()
    # 前后各保留50个字符,可按需调整,用max/min避免索引越界
    context_len = 50
    fragment_start = max(0, num_start - context_len)
    fragment_end = min(len(raw_text), num_end + context_len)
    target_fragment = raw_text[fragment_start:fragment_end]
    print(target_fragment)
    # 输出结果:, and this is not a sentence with numbers because 123.

如果需要处理文本中存在多个数字的场景,把re.search替换为re.finditer遍历所有匹配结果即可。


内容的提问来源于stack exchange,提问作者Tab1e

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 03:45:03