You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfquery无法精准匹配PDF题号:如何解决模糊匹配问题

解决PDF文本精准匹配题号的问题

你遇到的问题是:contains()选择器的包含匹配特性导致的——1.30的文本里包含"1.3",所以会被优先命中。要实现精准匹配,需要直接对比文本内容的完全一致,而不是包含关系。

解决方案:使用filter精准匹配文本

修改代码,遍历所有LTTextBoxHorizontal元素,通过filter方法筛选出文本内容与目标题号完全一致的元素(记得去除文本前后的空白字符,避免因PDF排版的空格导致匹配失败):

def search_terms():
    for q in questions:
        # 筛选文本内容完全匹配目标题号的文本框
        matched_boxes = pdf.pq('LTTextBoxHorizontal').filter(lambda i, e: e.text_content().strip() == q)
        if matched_boxes:
            answer = matched_boxes[0]
            x0 = float(answer.get("x0", 0)) + 350
            y0 = float(answer.get("y0", 0)) 
            x1 = float(answer.get("x1", 0)) + 480
            y1 = float(answer.get("y1", 0)) 
            locations.append((x0, y0, x1, y1))

补充说明

如果题号文本后可能紧跟其他内容(比如题目开头的字符),但你只需要匹配题号部分,可以用正则表达式做精确匹配:

import re

def search_terms():
    for q in questions:
        # 用正则精确匹配题号(确保是独立的题号,比如1.3后面没有额外数字)
        pattern = re.compile(f'^{re.escape(q)}$')
        matched_boxes = pdf.pq('LTTextBoxHorizontal').filter(lambda i, e: pattern.match(e.text_content().strip()))
        if matched_boxes:
            answer = matched_boxes[0]
            x0 = float(answer.get("x0", 0)) + 350
            y0 = float(answer.get("y0", 0)) 
            x1 = float(answer.get("x1", 0)) + 480
            y1 = float(answer.get("y1", 0)) 
            locations.append((x0, y0, x1, y1))

这样就能彻底避免"1.3"匹配到"1.30"的问题,确保只选中目标题号对应的文本框。

内容的提问来源于stack exchange,提问作者Genysgwav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 07:15:41