如何用Python提取PDF文本且不包含批注内容?
解决PyMuPDF提取PDF文本时排除批注内容且不遗漏原文本的问题
问题根源
你之前的代码通过矩形交集排除所有与批注重叠的单词,这会误删批注下方的原页面文本。核心问题在于没有区分批注自身的文本和批注覆盖的原文本:
- 高亮、下划线这类批注仅标记原文本,本身无额外文本,不需要排除原文本;
- 文本框、注释笔记这类批注自带独立文本,才需要排除这些额外文本。
解决方案
- 分类处理批注:仅针对包含独立文本的批注(如自由文本、文本注释)进行过滤;
- 提取批注自身的文本内容,而非通过矩形粗暴排除;
- 对自由文本批注,通过其文本所在矩形排除对应单词,避免误删原文本。
修改后的代码
def get_annotation_texts_and_rects(page): """收集所有需要排除的批注文本和自由文本批注的矩形""" excluded_texts = set() excluded_rects = [] if page.annots(): for annot in page.annots(): # 处理文本注释(弹出式笔记),提取其内容 if annot.type[0] == 5: # Text annotation type if annot.contents: excluded_texts.update(annot.contents.split()) # 处理自由文本批注(文本框),提取其文本和显示矩形 elif annot.type[0] == 3: # FreeText annotation type if annot.contents: excluded_texts.update(annot.contents.split()) excluded_rects.append(annot.rect) # 高亮、下划线等无独立文本的批注,无需处理 return excluded_texts, excluded_rects def extract_words_from_box(page, rect, excluded_texts, excluded_rects): words = page.get_text("words") # 提取页面所有单词 words_in_box = [word for word in words if rect.intersects(fitz.Rect(word[:4]))] # 过滤条件:既不在排除矩形内,单词文本也不属于批注内容 filtered_words = [] for word in words_in_box: word_rect = fitz.Rect(word[:4]) # 检查是否在自由文本批注的矩形内(完全包含,避免误判重叠的原文本) in_excluded_rect = any(rect.contains(word_rect) for rect in excluded_rects) # 检查单词是否属于批注文本 is_excluded_text = word[4] in excluded_texts if not in_excluded_rect and not is_excluded_text: filtered_words.append(word) return filtered_words def print_text_in_boxes(pdf_path): material_box_def = (40, 85, 70, 705) # 物料编号区域 po_number_box_def = (330, 40, 570, 115) # PO编号区域 destination_box_def = (40, 200, 330, 300) # 目的地区域 other_boxes_definitions = { 'product_code': (72, -2, 131, 15), 'due_date': (217, -2, 270, 15), 'qty': (286, -2, 314, 15), 'net_price': (350, -2, 399, 15), 'material_revision_box': (80, 15, 192, 80) } po_number = None destination = "Unknown" extracted_data = [] try: doc = fitz.open(pdf_path) for page_num in range(len(doc)): page = doc[page_num] # 获取需要排除的批注文本和矩形 excluded_texts, excluded_rects = get_annotation_texts_and_rects(page) # 仅从第一页提取PO编号和目的地 if page_num == 0: po_number_rect = fitz.Rect(*po_number_box_def) words_in_po_number_box = extract_words_from_box(page, po_number_rect, excluded_texts, excluded_rects) po_numbers = [word[4] for word in words_in_po_number_box if word[4].isdigit() and len(word[4]) == 10] if po_numbers: po_number = po_numbers[0] destination_rect = fitz.Rect(*destination_box_def) words_in_destination_box = extract_words_from_box(page, destination_rect, excluded_texts, excluded_rects) destination_text = ' '.join(word[4] for word in words_in_destination_box) destination = determine_destination(destination_text) material_rect = fitz.Rect(*material_box_def) words_in_material_box = extract_words_from_box(page, material_rect, excluded_texts, excluded_rects) item_numbers = [word for word in words_in_material_box if word[4].isdigit() and len(word[4]) == 5] for item_number in item_numbers: item_number_rect = fitz.Rect(item_number[:4]) # 注意:需同步修改extract_item_info函数,传入excluded_texts和excluded_rects代替原annotations参数 item_info = extract_item_info(page, item_number_rect, other_boxes_definitions, po_number, destination, excluded_texts, excluded_rects) # 拆分产品代码和名称 if 'product_code' in item_info: description_parts = item_info['product_code'].split(' ', 1) if len(description_parts) == 2: item_info['product_code'] = description_parts[0] item_info['product_name'] = description_parts[1] else: item_info['product_code'] = description_parts[0] item_info['product_name'] = 'No text found' # 转换日期为芬兰格式 if 'due_date' in item_info: item_info['due_date'] = convert_date_format(item_info['due_date']) extracted_data.append(item_info) except Exception as e: print(f"处理{pdf_path}时出错: {str(e)}") return extracted_data
关键调整说明
- 新增
get_annotation_texts_and_rects函数:只收集需要排除的批注文本(文本注释、自由文本)和自由文本的显示矩形,高亮/下划线等批注不处理; - 修改
extract_words_from_box函数:- 用
rect.contains(word_rect)代替intersects,确保只排除完全在自由文本矩形内的单词(避免误删重叠的原文本); - 增加文本匹配过滤,排除属于批注内容的单词;
- 用
- 传递过滤参数:将
excluded_texts和excluded_rects传入所有提取文本的函数,替换原有的annotations参数(需同步修改extract_item_info函数的参数和内部逻辑)。
内容的提问来源于stack exchange,提问作者MDMT
相关产品推荐
相关产品推荐

