You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用fitz page.add_highlight_annot高亮PDF文本异常求助

问题根源分析

你当前的核心问题是对page.add_highlight_annot(start=pointa, stop=pointb)的参数理解错误:这里的start和stop并非页面坐标点,而是页面文本的字符索引位置(整数类型)。传入坐标点后fitz无法正确解析,因此默认高亮了整页内容。

另外你的类写法不规范——初始化逻辑直接放在类顶层会导致类定义时就执行代码,应放到__init__方法中。

两种可行修正方案

方案一:基于文本字符索引高亮

先获取页面完整文本,定位起始词和结束词的字符位置,再用这些索引创建高亮(适合排版简单的PDF):

import fitz
import config

class FindTextblock:
    def __init__(self):
        doc = fitz.open(config.fpath_int + "/test_saurce.pdf")
        for page in doc:
            page_text = page.get_text()
            # 定位起始词的起始索引
            start_idx = page_text.find("Therapievorschlag")
            # 定位结束词的结束索引(包含结束词本身)
            end_idx = page_text.find("EU-Verordnung") + len("EU-Verordnung")
            
            if start_idx != -1 and end_idx != -1:
                page.add_highlight_annot(start=start_idx, stop=end_idx)
        
        doc.save("test.pdf")

FindTextblock()

方案二:基于单词矩形合并高亮

针对排版复杂(如分栏、文字换行拆分)的PDF,收集起始词到结束词之间所有单词的矩形,合并后创建高亮:

import fitz
import config

class FindTextblock:
    def __init__(self):
        doc = fitz.open(config.fpath_int + "/test_saurce.pdf")
        for page in doc:
            wordlist = page.get_text("words")
            # 按垂直、水平排序,保证阅读顺序正确
            wordlist.sort(key=lambda w: (w[1], w[0]))
            start_found = False
            highlight_rects = []
            
            for w in wordlist:
                word_text = w[4]
                if word_text == "Therapievorschlag":
                    start_found = True
                    highlight_rects.append(fitz.Rect(w[:4]))
                elif start_found and word_text == "EU-Verordnung":
                    highlight_rects.append(fitz.Rect(w[:4]))
                    break
                elif start_found:
                    highlight_rects.append(fitz.Rect(w[:4]))
            
            if highlight_rects:
                # 合并所有矩形
                combined_rect = fitz.Rect()
                for rect in highlight_rects:
                    combined_rect |= rect
                page.add_highlight_annot(combined_rect)
        
        doc.save("test.pdf")

FindTextblock()
注意事项
  • 方案一中find方法只会返回第一个匹配的位置,若页面存在多个相同关键词,需额外处理匹配逻辑。
  • 方案二更适配复杂排版场景,基于单词矩形的高亮精度更高。

内容的提问来源于stack exchange,提问作者kalimero00

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 08:47:20