如何用正则提取引用上下文?解决索引越界问题
问题:提取引用及对应上下文时触发IndexError
我的代码用于从文本中提取引用/参考文献,以及引用左右最多10个字符的上下文:
import re # some toy text text = 'Once upon a time a cat says «gross!». A long story you can check here (ref. 11). People witnessed the scene [...]' quoting_pattern = '\([^\(]*\)|„[^„]*"|<<.*>>|«[^«]*»|“[^“]*”|‹[^‹]*›|"[^"]*"|›[^›]*‹|»[^»]*«' context_pattern = ".{0,100}(?:{}).{0,100}".format(quoting_pattern) # get all quotations quotations = re.findall(r'{}'.format(quoting_pattern), text, re.DOTALL) # get all contexts contexts = re.findall(r'{}'.format(context_pattern), text, re.DOTALL) for i, q in enumerate(quotations): print(q, contexts[i])
预期输出:
"«gross!»", " cat says «gross!». A long s" "(ref. 11)", "heck here (ref. 11)"
但运行时出现IndexError: list index out of range:quotations能提取到«gross!»和(ref. 11),但contexts只有前者的上下文,后者无法匹配到。
问题原因
- 正则贪婪匹配导致覆盖:
context_pattern中的.{0,100}是贪婪匹配模式,第一个上下文匹配会尽可能向后延伸,吃掉第二个引用所在的文本内容,导致re.findall只能找到1个上下文结果,而quotations有2个元素,索引越界。 - 上下文长度定义不符:代码中用了
.{0,100},但需求是“左右最多10个字符”,长度范围错误。
解决方法
方法1:修正正则为非贪婪匹配,调整长度范围
将context_pattern中的贪婪匹配改为非贪婪匹配(在量词后加?),同时把长度从100改为10,符合需求:
import re text = 'Once upon a time a cat says «gross!». A long story you can check here (ref. 11). People witnessed the scene [...]' quoting_pattern = r'\([^\(]*\)|„[^„]*"|<<.*>>|«[^«]*»|“[^“]*”|‹[^‹]*›|"[^"]*"|›[^›]*‹|»[^»]*«' # 改为非贪婪匹配,且长度调整为0-10 context_pattern = r".{0,10}?(?:{}).{0,10}?".format(quoting_pattern) quotations = re.findall(quoting_pattern, text, re.DOTALL) contexts = re.findall(context_pattern, text, re.DOTALL) for i, q in enumerate(quotations): print(f'"{q}", "{contexts[i]}"')
方法2:使用捕获组同时提取上下文和引用(更可靠)
通过一个正则同时匹配引用的前后上下文和引用本身,确保每个引用都能对应到上下文,避免数量不一致:
import re text = 'Once upon a time a cat says «gross!». A long story you can check here (ref. 11). People witnessed the scene [...]' quoting_pattern = r'\([^\(]*\)|„[^„]*"|<<.*>>|«[^«]*»|“[^“]*”|‹[^‹]*›|"[^"]*"|›[^›]*‹|»[^»]*«' # 匹配引用前0-10字符、引用本身、引用后0-10字符 combined_pattern = r'(?P<before>.{0,10})(?P<quote>{quoting_pattern})(?P<after>.{0,10})'.format(quoting_pattern=quoting_pattern) matches = re.finditer(combined_pattern, text, re.DOTALL) for match in matches: quote = match.group('quote') context = match.group('before') + quote + match.group('after') print(f'"{quote}", "{context}"')
运行后可得到预期输出:
"«gross!»", " cat says «gross!». A long s" "(ref. 11)", "heck here (ref. 11)"
内容的提问来源于stack exchange,提问作者Ding Dong
相关产品推荐
相关产品推荐

