如何用Python(PyMuPDF)实现PDF替换文本与原文本排版、字体及大小匹配
问题描述
使用Python和PyMuPDF库进行PDF文本搜索替换时,替换后的文本字体、大小与原文本不一致,位置也无法精准匹配。空白替换文本可正常清除原文本,但输入替换内容时就会出现上述问题。原代码如下:
import os import fitz # Prompt user for input of file name file_name_input = input("Enter a start or text of the file name: ") # Get a list of PDF files in the current directory matching the file name input pdf_files = [f for f in os.listdir() if f.lower().endswith('.pdf') and file_name_input.lower() in f.lower()] if not pdf_files: print("No PDF files found matching the file name input") else: # Prompt user for input of search and replace text search_replace_list = [] while True: search_text = input("Enter the search text (leave blank to exit): ") if not search_text: break replace_text = input("Enter the replace text: ") search_replace_list.append((search_text, replace_text)) for file_name in pdf_files: pdf_file = fitz.open(file_name) found = False for page in pdf_file: for search_text, replace_text in search_replace_list: draft = page.search_for(search_text.strip(), hit_max=16, quads=True, quads_tol=0.01) if draft: found = True for rect in draft: annot = page.add_redact_annot(rect, text=replace_text) page.apply_redactions() page.apply_redactions(images=fitz.PDF_REDACT_IMAGE_NONE) if found: output_file_name = file_name[:-4] + '_modified.pdf' pdf_file.save(output_file_name, garbage=False, deflate=True, encryption=False) print(f"Changes saved to {output_file_name}") else: print(f"No search text found in {file_name}") pdf_file.close()
解决方案
原代码使用add_redact_annot时,默认会用系统基础字体插入替换文本,导致样式不匹配。要实现样式和位置完全一致,需按以下步骤处理:
- 搜索到目标文本的位置后,通过
page.get_text("words")获取页面所有文本块的详细信息(包含字体、字号、颜色、位置矩形) - 匹配搜索到的位置与文本块,提取原文本的样式属性
- 先通过红act操作清除原文本,再用
insert_textbox以原样式插入替换文本,确保位置和排版一致 - 空白替换时直接执行红act清除即可
修改后的代码
import os import fitz def get_text_style(page, search_rect): """获取指定矩形区域内文本的样式属性(字体、字号、颜色)""" words = page.get_text("words", flags=fitz.TEXTFLAGS_TEXT) for word in words: word_rect = fitz.Rect(word[:4]) # 判断当前文本块是否与搜索矩形重叠(容错范围) if word_rect.intersects(search_rect) and word_rect.get_area() * 0.8 < search_rect.get_area(): # 提取字体信息 font_info = page.get_font_info(word[4]) font_name = font_info["fontname"] # 字号取文本块的高度近似值 font_size = word_rect.height # 文本颜色(默认黑色) color = word[5] if len(word) >5 else (0,0,0) return font_name, font_size, color # 默认返回系统字体 return ("helv", 12, (0,0,0)) # Prompt user for input of file name file_name_input = input("Enter a start or text of the file name: ") # Get a list of PDF files in the current directory matching the file name input pdf_files = [f for f in os.listdir() if f.lower().endswith('.pdf') and file_name_input.lower() in f.lower()] if not pdf_files: print("No PDF files found matching the file name input") else: # Prompt user for input of search and replace text search_replace_list = [] while True: search_text = input("Enter the search text (leave blank to exit): ") if not search_text: break replace_text = input("Enter the replace text: ") search_replace_list.append((search_text.strip(), replace_text)) for file_name in pdf_files: pdf_file = fitz.open(file_name) found = False for page in pdf_file: for search_text, replace_text in search_replace_list: # 搜索目标文本,获取位置矩形 hits = page.search_for(search_text, hit_max=16, quads=False) if hits: found = True for hit_rect in hits: # 第一步:添加红act注释清除原文本 page.add_redact_annot(hit_rect) # 应用红act清除 page.apply_redactions(images=fitz.PDF_REDACT_IMAGE_NONE) # 如果替换文本不为空,插入匹配样式的新文本 if replace_text: for hit_rect in hits: # 获取原文本样式 font_name, font_size, color = get_text_style(page, hit_rect) # 将颜色转换为PyMuPDF支持的RGB格式 rgb_color = fitz.sRGB_to_pdf(color) # 插入文本,使用原字体和字号,居中对齐匹配原位置 page.insert_textbox( hit_rect, replace_text, fontname=font_name, fontsize=font_size, color=rgb_color, align=fitz.TEXT_ALIGN_CENTER, valign=fitz.TEXT_ALIGN_CENTER ) if found: output_file_name = file_name[:-4] + '_modified.pdf' # 保存时保留原文档的字体资源 pdf_file.save(output_file_name, garbage=3, deflate=True, encryption=False) print(f"Changes saved to {output_file_name}") else: print(f"No search text found in {file_name}") pdf_file.close()
关键说明
get_text_style函数:通过页面文本块信息匹配搜索区域,提取原文本的字体名称、字号和颜色- 红act操作仅用于清除原文本,替换文本通过
insert_textbox插入,确保使用原样式 - 插入文本时设置居中对齐,保证位置与原文本完全匹配
- 保存文档时使用
garbage=3,确保保留原文档的字体资源,避免字体缺失
内容的提问来源于stack exchange,提问作者Sik Saw
相关产品推荐
相关产品推荐

