使用PDFminer.six提取PDF文本后,如何捕获特定短语及周边文本?
解决方法
要找到短语的所有出现并保存周边文本,最简便的方式是用正则表达式定位所有匹配位置,再根据位置截取上下文内容,具体实现如下:
完整代码示例
from pdfminer.high_level import extract_text import re # 提取PDF文本 text = extract_text('Pdf Scanner/test.pdf') # 目标短语 target_phrase = "vejkode" # 定义上下文长度(前后各取n个字符,可按需调整) context_length = 50 # 存储匹配结果的列表 matches_with_context = [] # 遍历所有匹配项 for match in re.finditer(re.escape(target_phrase), text): # 计算上下文起始索引(避免小于0) start_idx = max(0, match.start() - context_length) # 计算上下文结束索引(避免超过文本总长度) end_idx = min(len(text), match.end() + context_length) # 截取周边文本 context = text[start_idx:end_idx] # 保存匹配详情 matches_with_context.append({ 'position': match.start(), 'phrase': match.group(), 'context': context }) # 查看结果 for idx, item in enumerate(matches_with_context, 1): print(f"匹配项 {idx}:") print(f"位置:{item['position']}") print(f"上下文:{item['context']}\n")
关键说明
re.escape(target_phrase):如果目标短语包含正则特殊字符(如.、*),用该方法转义可避免匹配逻辑出错- 索引边界处理:通过
max(0, ...)和min(len(text), ...)确保不会因上下文长度设置过大导致索引越界 - 结果存储:用字典保存每个匹配的位置、短语和上下文,后续可直接遍历列表调用所需内容
内容的提问来源于stack exchange,提问作者SykesTheLord
相关产品推荐
相关产品推荐

