如何用Python+PyMuPDF实现PDF替换文本与原文本的格式样式匹配?
PyMuPDF替换PDF文本保留原样式的实现方法
我用Python的PyMuPDF库做PDF文本搜索替换,功能能正常运行,但替换后的文本没法匹配原文本的颜色、字号等样式,求解决办法。当前使用的代码如下:
import os import fitz # Prompt user for input of file name file_name_input = input("Enter a start or text of the file name: ") # Get a list of PDF files in the current directory matching the file name input pdf_files = [f for f in os.listdir() if f.lower().endswith('.pdf') and file_name_input.lower() in f.lower()] if not pdf_files: print("No PDF files found matching the file name input") else: # Prompt user for input of search and replace text search_replace_list = [] while True: search_text = input("Enter the search text (leave blank to exit): ") if not search_text: break replace_text = input("Enter the replace text: ") search_replace_list.append((search_text, replace_text)) for file_name in pdf_files: pdf_file = fitz.open(file_name) found = False for page in pdf_file: for search_text, replace_text in search_replace_list: draft = page.search_for(search_text.strip(), hit_max=16, quads=True, quads_tol=0.01) if draft: found = True for rect in draft: annot = page.add_redact_annot(rect, text=replace_text) page.apply_redactions() page.apply_redactions(images=fitz.PDF_REDACT_IMAGE_NONE) if found: output_file_name = file_name[:-4] + '_modified.pdf' pdf_file.save(output_file_name, garbage=False, deflate=True, encryption=False) print(f"Changes saved to {output_file_name}") else: print(f"No search text found in {file_name}") pdf_file.close()
问题原因
当前代码使用add_redact_annot进行替换时,默认采用PyMuPDF的内置文本样式,没有继承原文本的字体、颜色、字号等属性,导致替换后的样式和原文本不一致。
解决思路
要保留原样式,需要先获取待替换文本的原始样式信息(字体名称、字号、颜色等),再在替换时应用这些样式:
- 遍历页面文本块,匹配搜索文本并提取对应样式属性
- 清除原文本所在区域内容
- 使用原始样式插入替换文本
修改后的代码
import os import fitz # 匹配文本并获取其样式信息 def find_text_with_style(page, search_text): text_dict = page.get_text("dict") matches = [] for block in text_dict["blocks"]: if "lines" not in block: continue for line in block["lines"]: for span in line["spans"]: if search_text in span["text"]: # 计算匹配文本的位置矩形(系数0.6为字体宽度估算值,可按需调整) char_width = span["size"] * 0.6 start_x = span["bbox"][0] + span["text"].index(search_text) * char_width rect = fitz.Rect( start_x, span["bbox"][1], start_x + len(search_text) * char_width, span["bbox"][3] ) matches.append({ "rect": rect, "font": span["font"], "size": span["size"], "color": span["color"], "flags": span["flags"] }) return matches # 主逻辑 file_name_input = input("Enter a start or text of the file name: ") pdf_files = [f for f in os.listdir() if f.lower().endswith('.pdf') and file_name_input.lower() in f.lower()] if not pdf_files: print("No PDF files found matching the file name input") else: search_replace_list = [] while True: search_text = input("Enter the search text (leave blank to exit): ") if not search_text: break replace_text = input("Enter the replace text: ") search_replace_list.append((search_text, replace_text)) for file_name in pdf_files: pdf_file = fitz.open(file_name) found = False for page in pdf_file: for search_text, replace_text in search_replace_list: matches = find_text_with_style(page, search_text) if matches: found = True # 先清除原文本区域 for match in matches: page.add_redact_annot(match["rect"]) page.apply_redactions() # 插入带原样式的替换文本 for match in matches: # 转换颜色格式为PyMuPDF可用格式 color = fitz.sRGB_to_pdf(match["color"]) page.insert_text( match["rect"].tl, # 以原文本左上角为插入起点 replace_text, fontname=match["font"], fontsize=match["size"], color=color, flags=match["flags"] ) if found: output_file_name = file_name[:-4] + '_modified.pdf' pdf_file.save(output_file_name, garbage=False, deflate=True, encryption=False) print(f"Changes saved to {output_file_name}") else: print(f"No search text found in {file_name}") pdf_file.close()
注意事项
- 文本位置计算:使用字号×0.6估算字符宽度,不同字体可能存在偏差,可根据实际情况调整系数
- 字体兼容性:若PDF使用嵌入字体,PyMuPDF需能识别该字体才能正确应用样式,否则会 fallback 到默认字体
- 多行文本匹配:当前代码仅支持单行内的文本匹配,若搜索文本跨多行,需额外处理文本块的拼接逻辑
内容的提问来源于stack exchange,提问作者Hetul
相关产品推荐
相关产品推荐

