使用fitz page.add_highlight_annot高亮PDF文本异常求助
问题根源分析
你当前的核心问题是对page.add_highlight_annot(start=pointa, stop=pointb)的参数理解错误:这里的start和stop并非页面坐标点,而是页面文本的字符索引位置(整数类型)。传入坐标点后fitz无法正确解析,因此默认高亮了整页内容。
另外你的类写法不规范——初始化逻辑直接放在类顶层会导致类定义时就执行代码,应放到__init__方法中。
两种可行修正方案
方案一:基于文本字符索引高亮
先获取页面完整文本,定位起始词和结束词的字符位置,再用这些索引创建高亮(适合排版简单的PDF):
import fitz import config class FindTextblock: def __init__(self): doc = fitz.open(config.fpath_int + "/test_saurce.pdf") for page in doc: page_text = page.get_text() # 定位起始词的起始索引 start_idx = page_text.find("Therapievorschlag") # 定位结束词的结束索引(包含结束词本身) end_idx = page_text.find("EU-Verordnung") + len("EU-Verordnung") if start_idx != -1 and end_idx != -1: page.add_highlight_annot(start=start_idx, stop=end_idx) doc.save("test.pdf") FindTextblock()
方案二:基于单词矩形合并高亮
针对排版复杂(如分栏、文字换行拆分)的PDF,收集起始词到结束词之间所有单词的矩形,合并后创建高亮:
import fitz import config class FindTextblock: def __init__(self): doc = fitz.open(config.fpath_int + "/test_saurce.pdf") for page in doc: wordlist = page.get_text("words") # 按垂直、水平排序,保证阅读顺序正确 wordlist.sort(key=lambda w: (w[1], w[0])) start_found = False highlight_rects = [] for w in wordlist: word_text = w[4] if word_text == "Therapievorschlag": start_found = True highlight_rects.append(fitz.Rect(w[:4])) elif start_found and word_text == "EU-Verordnung": highlight_rects.append(fitz.Rect(w[:4])) break elif start_found: highlight_rects.append(fitz.Rect(w[:4])) if highlight_rects: # 合并所有矩形 combined_rect = fitz.Rect() for rect in highlight_rects: combined_rect |= rect page.add_highlight_annot(combined_rect) doc.save("test.pdf") FindTextblock()
注意事项
- 方案一中
find方法只会返回第一个匹配的位置,若页面存在多个相同关键词,需额外处理匹配逻辑。 - 方案二更适配复杂排版场景,基于单词矩形的高亮精度更高。
内容的提问来源于stack exchange,提问作者kalimero00
相关产品推荐
相关产品推荐

