如何截取re.findall匹配结果的前置文本?解决TypeError问题
问题描述
原始文本:
some text some text Jack is the CEO. some text some text John DOE is the CEO.
编写的目标函数:
def get_ceo(text): results = re.findall(r"is the CEO", text) for i in results: range = text[i-15:i] print(range)
期望输出:
['some text Jack is the CEO',' text John DOE is the CEO']
运行时抛出错误:
line 62, in <module> print(get_ceo(text)) line 50, in get_ceo range = text[i-15:i] TypeError: unsupported operand type(s) for -: 'str' and 'int'
疑问:是否需要转换re.findall结果的类型,或是需要完全更换实现思路?
解决方案
错误原因很直接:re.findall返回的是匹配到的字符串列表(每个元素都是"is the CEO"),你在循环里把字符串当作下标去做减法操作,必然触发类型错误。
不需要转换类型,直接更换实现思路即可——用re.finditer替代re.findall,它会返回包含匹配位置信息的Match对象,能获取每个匹配结果的起始/结束索引:
import re def get_ceo(text): pattern = r"is the CEO" ceo_segments = [] for match in re.finditer(pattern, text): # 计算截取起始位置,避免负索引越界 start_pos = max(0, match.start() - 15) # 截取从起始位置到匹配结束的完整片段(包含匹配内容) segment = text[start_pos:match.end()] ceo_segments.append(segment) return ceo_segments # 测试示例 text = "some text some text Jack is the CEO. some text some text John DOE is the CEO. " print(get_ceo(text))
运行后会得到期望的输出:
['some text Jack is the CEO', ' text John DOE is the CEO']
关键说明
match.start():获取匹配字符串在原文本中的起始下标match.end():获取匹配字符串在原文本中的结束下标max(0, ...):处理边界情况,当匹配结果前不足15个字符时,从文本开头开始截取,避免负索引报错
内容的提问来源于stack exchange,提问作者user16779293
相关产品推荐
相关产品推荐

