You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则提取引用上下文?解决索引越界问题

问题:提取引用及对应上下文时触发IndexError

我的代码用于从文本中提取引用/参考文献,以及引用左右最多10个字符的上下文:

import re

# some toy text
text = 'Once upon a time a cat says «gross!». A long story you can check here (ref. 11). People witnessed the scene [...]'

quoting_pattern = '\([^\(]*\)|„[^„]*"|<<.*>>|«[^«]*»|“[^“]*”|‹[^‹]*›|"[^"]*"|›[^›]*‹|»[^»]*«'
context_pattern = ".{0,100}(?:{}).{0,100}".format(quoting_pattern)

# get all quotations
quotations = re.findall(r'{}'.format(quoting_pattern), text, re.DOTALL)
    
# get all contexts
contexts = re.findall(r'{}'.format(context_pattern), text, re.DOTALL)

for i, q in enumerate(quotations):
    print(q, contexts[i])

预期输出:

"«gross!»", " cat says «gross!». A long s"
"(ref. 11)", "heck here (ref. 11)"

但运行时出现IndexError: list index out of range:quotations能提取到«gross!»和(ref. 11),但contexts只有前者的上下文,后者无法匹配到。


问题原因

  1. 正则贪婪匹配导致覆盖:context_pattern中的.{0,100}是贪婪匹配模式,第一个上下文匹配会尽可能向后延伸,吃掉第二个引用所在的文本内容,导致re.findall只能找到1个上下文结果,而quotations有2个元素,索引越界。
  2. 上下文长度定义不符:代码中用了.{0,100},但需求是“左右最多10个字符”,长度范围错误。

解决方法

方法1:修正正则为非贪婪匹配,调整长度范围

将context_pattern中的贪婪匹配改为非贪婪匹配(在量词后加?),同时把长度从100改为10,符合需求:

import re

text = 'Once upon a time a cat says «gross!». A long story you can check here (ref. 11). People witnessed the scene [...]'

quoting_pattern = r'\([^\(]*\)|„[^„]*"|<<.*>>|«[^«]*»|“[^“]*”|‹[^‹]*›|"[^"]*"|›[^›]*‹|»[^»]*«'
# 改为非贪婪匹配,且长度调整为0-10
context_pattern = r".{0,10}?(?:{}).{0,10}?".format(quoting_pattern)

quotations = re.findall(quoting_pattern, text, re.DOTALL)
contexts = re.findall(context_pattern, text, re.DOTALL)

for i, q in enumerate(quotations):
    print(f'"{q}", "{contexts[i]}"')

方法2:使用捕获组同时提取上下文和引用(更可靠)

通过一个正则同时匹配引用的前后上下文和引用本身,确保每个引用都能对应到上下文,避免数量不一致:

import re

text = 'Once upon a time a cat says «gross!». A long story you can check here (ref. 11). People witnessed the scene [...]'

quoting_pattern = r'\([^\(]*\)|„[^„]*"|<<.*>>|«[^«]*»|“[^“]*”|‹[^‹]*›|"[^"]*"|›[^›]*‹|»[^»]*«'
# 匹配引用前0-10字符、引用本身、引用后0-10字符
combined_pattern = r'(?P<before>.{0,10})(?P<quote>{quoting_pattern})(?P<after>.{0,10})'.format(quoting_pattern=quoting_pattern)

matches = re.finditer(combined_pattern, text, re.DOTALL)

for match in matches:
    quote = match.group('quote')
    context = match.group('before') + quote + match.group('after')
    print(f'"{quote}", "{context}"')

运行后可得到预期输出:

"«gross!»", " cat says «gross!». A long s"
"(ref. 11)", "heck here (ref. 11)"

内容的提问来源于stack exchange,提问作者Ding Dong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 06:06:20