You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取PDF文本且不包含批注内容?

解决PyMuPDF提取PDF文本时排除批注内容且不遗漏原文本的问题

问题根源

你之前的代码通过矩形交集排除所有与批注重叠的单词,这会误删批注下方的原页面文本。核心问题在于没有区分批注自身的文本和批注覆盖的原文本:

  • 高亮、下划线这类批注仅标记原文本,本身无额外文本,不需要排除原文本;
  • 文本框、注释笔记这类批注自带独立文本,才需要排除这些额外文本。

解决方案

  1. 分类处理批注:仅针对包含独立文本的批注(如自由文本、文本注释)进行过滤;
  2. 提取批注自身的文本内容,而非通过矩形粗暴排除;
  3. 对自由文本批注,通过其文本所在矩形排除对应单词,避免误删原文本。

修改后的代码

def get_annotation_texts_and_rects(page):
    """收集所有需要排除的批注文本和自由文本批注的矩形"""
    excluded_texts = set()
    excluded_rects = []
    if page.annots():
        for annot in page.annots():
            # 处理文本注释(弹出式笔记),提取其内容
            if annot.type[0] == 5:  # Text annotation type
                if annot.contents:
                    excluded_texts.update(annot.contents.split())
            # 处理自由文本批注(文本框),提取其文本和显示矩形
            elif annot.type[0] == 3:  # FreeText annotation type
                if annot.contents:
                    excluded_texts.update(annot.contents.split())
                    excluded_rects.append(annot.rect)
            # 高亮、下划线等无独立文本的批注,无需处理
    return excluded_texts, excluded_rects

def extract_words_from_box(page, rect, excluded_texts, excluded_rects):
    words = page.get_text("words")  # 提取页面所有单词
    words_in_box = [word for word in words if rect.intersects(fitz.Rect(word[:4]))]
    
    # 过滤条件:既不在排除矩形内,单词文本也不属于批注内容
    filtered_words = []
    for word in words_in_box:
        word_rect = fitz.Rect(word[:4])
        # 检查是否在自由文本批注的矩形内(完全包含,避免误判重叠的原文本)
        in_excluded_rect = any(rect.contains(word_rect) for rect in excluded_rects)
        # 检查单词是否属于批注文本
        is_excluded_text = word[4] in excluded_texts
        if not in_excluded_rect and not is_excluded_text:
            filtered_words.append(word)
    return filtered_words

def print_text_in_boxes(pdf_path):
    material_box_def = (40, 85, 70, 705)  # 物料编号区域
    po_number_box_def = (330, 40, 570, 115)  # PO编号区域
    destination_box_def = (40, 200, 330, 300)  # 目的地区域
    other_boxes_definitions = {
        'product_code': (72, -2, 131, 15),
        'due_date': (217, -2, 270, 15),
        'qty': (286, -2, 314, 15),
        'net_price': (350, -2, 399, 15),
        'material_revision_box': (80, 15, 192, 80)
    }

    po_number = None
    destination = "Unknown"
    extracted_data = []

    try:
        doc = fitz.open(pdf_path)

        for page_num in range(len(doc)):
            page = doc[page_num]
            # 获取需要排除的批注文本和矩形
            excluded_texts, excluded_rects = get_annotation_texts_and_rects(page)

            # 仅从第一页提取PO编号和目的地
            if page_num == 0:
                po_number_rect = fitz.Rect(*po_number_box_def)
                words_in_po_number_box = extract_words_from_box(page, po_number_rect, excluded_texts, excluded_rects)
                po_numbers = [word[4] for word in words_in_po_number_box if word[4].isdigit() and len(word[4]) == 10]
                if po_numbers:
                    po_number = po_numbers[0]

                destination_rect = fitz.Rect(*destination_box_def)
                words_in_destination_box = extract_words_from_box(page, destination_rect, excluded_texts, excluded_rects)
                destination_text = ' '.join(word[4] for word in words_in_destination_box)
                destination = determine_destination(destination_text)

            material_rect = fitz.Rect(*material_box_def)
            words_in_material_box = extract_words_from_box(page, material_rect, excluded_texts, excluded_rects)
            
            item_numbers = [word for word in words_in_material_box if word[4].isdigit() and len(word[4]) == 5]
            for item_number in item_numbers:
                item_number_rect = fitz.Rect(item_number[:4])
                # 注意:需同步修改extract_item_info函数,传入excluded_texts和excluded_rects代替原annotations参数
                item_info = extract_item_info(page, item_number_rect, other_boxes_definitions, po_number, destination, excluded_texts, excluded_rects)
                
                # 拆分产品代码和名称
                if 'product_code' in item_info:
                    description_parts = item_info['product_code'].split(' ', 1)
                    if len(description_parts) == 2:
                        item_info['product_code'] = description_parts[0]
                        item_info['product_name'] = description_parts[1]
                    else:
                        item_info['product_code'] = description_parts[0]
                        item_info['product_name'] = 'No text found'

                # 转换日期为芬兰格式
                if 'due_date' in item_info:
                    item_info['due_date'] = convert_date_format(item_info['due_date'])

                extracted_data.append(item_info)

    except Exception as e:
        print(f"处理{pdf_path}时出错: {str(e)}")

    return extracted_data

关键调整说明

  • 新增get_annotation_texts_and_rects函数:只收集需要排除的批注文本(文本注释、自由文本)和自由文本的显示矩形,高亮/下划线等批注不处理;
  • 修改extract_words_from_box函数:
    • 用rect.contains(word_rect)代替intersects,确保只排除完全在自由文本矩形内的单词(避免误删重叠的原文本);
    • 增加文本匹配过滤,排除属于批注内容的单词;
  • 传递过滤参数:将excluded_texts和excluded_rects传入所有提取文本的函数,替换原有的annotations参数(需同步修改extract_item_info函数的参数和内部逻辑)。

内容的提问来源于stack exchange,提问作者MDMT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 12:24:57