You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python(PyMuPDF)实现PDF替换文本与原文本排版、字体及大小匹配

问题描述

使用Python和PyMuPDF库进行PDF文本搜索替换时,替换后的文本字体、大小与原文本不一致,位置也无法精准匹配。空白替换文本可正常清除原文本,但输入替换内容时就会出现上述问题。原代码如下:

import os
import fitz

# Prompt user for input of file name
file_name_input = input("Enter a start or text of the file name: ")

# Get a list of PDF files in the current directory matching the file name input
pdf_files = [f for f in os.listdir() if f.lower().endswith('.pdf') and file_name_input.lower() in f.lower()]

if not pdf_files:
    print("No PDF files found matching the file name input")
else:
    # Prompt user for input of search and replace text
    search_replace_list = []
    while True:
        search_text = input("Enter the search text (leave blank to exit): ")
        if not search_text:
            break
        replace_text = input("Enter the replace text: ")
        search_replace_list.append((search_text, replace_text))

    for file_name in pdf_files:
        pdf_file = fitz.open(file_name)
        found = False
        for page in pdf_file:
            for search_text, replace_text in search_replace_list:
                draft = page.search_for(search_text.strip(), hit_max=16, quads=True, quads_tol=0.01)
                if draft:
                    found = True
                    for rect in draft:
                        annot = page.add_redact_annot(rect, text=replace_text)
                    page.apply_redactions()
                    page.apply_redactions(images=fitz.PDF_REDACT_IMAGE_NONE)

        if found:
            output_file_name = file_name[:-4] + '_modified.pdf'
            pdf_file.save(output_file_name, garbage=False, deflate=True, encryption=False)
            print(f"Changes saved to {output_file_name}")
        else:
            print(f"No search text found in {file_name}")

        pdf_file.close()
解决方案

原代码使用add_redact_annot时,默认会用系统基础字体插入替换文本,导致样式不匹配。要实现样式和位置完全一致,需按以下步骤处理:

  • 搜索到目标文本的位置后,通过page.get_text("words")获取页面所有文本块的详细信息(包含字体、字号、颜色、位置矩形)
  • 匹配搜索到的位置与文本块,提取原文本的样式属性
  • 先通过红act操作清除原文本,再用insert_textbox以原样式插入替换文本,确保位置和排版一致
  • 空白替换时直接执行红act清除即可
修改后的代码
import os
import fitz

def get_text_style(page, search_rect):
    """获取指定矩形区域内文本的样式属性(字体、字号、颜色)"""
    words = page.get_text("words", flags=fitz.TEXTFLAGS_TEXT)
    for word in words:
        word_rect = fitz.Rect(word[:4])
        # 判断当前文本块是否与搜索矩形重叠(容错范围)
        if word_rect.intersects(search_rect) and word_rect.get_area() * 0.8 < search_rect.get_area():
            # 提取字体信息
            font_info = page.get_font_info(word[4])
            font_name = font_info["fontname"]
            # 字号取文本块的高度近似值
            font_size = word_rect.height
            # 文本颜色(默认黑色)
            color = word[5] if len(word) >5 else (0,0,0)
            return font_name, font_size, color
    # 默认返回系统字体
    return ("helv", 12, (0,0,0))

# Prompt user for input of file name
file_name_input = input("Enter a start or text of the file name: ")

# Get a list of PDF files in the current directory matching the file name input
pdf_files = [f for f in os.listdir() if f.lower().endswith('.pdf') and file_name_input.lower() in f.lower()]

if not pdf_files:
    print("No PDF files found matching the file name input")
else:
    # Prompt user for input of search and replace text
    search_replace_list = []
    while True:
        search_text = input("Enter the search text (leave blank to exit): ")
        if not search_text:
            break
        replace_text = input("Enter the replace text: ")
        search_replace_list.append((search_text.strip(), replace_text))

    for file_name in pdf_files:
        pdf_file = fitz.open(file_name)
        found = False
        for page in pdf_file:
            for search_text, replace_text in search_replace_list:
                # 搜索目标文本,获取位置矩形
                hits = page.search_for(search_text, hit_max=16, quads=False)
                if hits:
                    found = True
                    for hit_rect in hits:
                        # 第一步:添加红act注释清除原文本
                        page.add_redact_annot(hit_rect)
                    # 应用红act清除
                    page.apply_redactions(images=fitz.PDF_REDACT_IMAGE_NONE)

                    # 如果替换文本不为空,插入匹配样式的新文本
                    if replace_text:
                        for hit_rect in hits:
                            # 获取原文本样式
                            font_name, font_size, color = get_text_style(page, hit_rect)
                            # 将颜色转换为PyMuPDF支持的RGB格式
                            rgb_color = fitz.sRGB_to_pdf(color)
                            # 插入文本,使用原字体和字号,居中对齐匹配原位置
                            page.insert_textbox(
                                hit_rect,
                                replace_text,
                                fontname=font_name,
                                fontsize=font_size,
                                color=rgb_color,
                                align=fitz.TEXT_ALIGN_CENTER,
                                valign=fitz.TEXT_ALIGN_CENTER
                            )

        if found:
            output_file_name = file_name[:-4] + '_modified.pdf'
            # 保存时保留原文档的字体资源
            pdf_file.save(output_file_name, garbage=3, deflate=True, encryption=False)
            print(f"Changes saved to {output_file_name}")
        else:
            print(f"No search text found in {file_name}")

        pdf_file.close()
关键说明
  • get_text_style函数:通过页面文本块信息匹配搜索区域,提取原文本的字体名称、字号和颜色
  • 红act操作仅用于清除原文本,替换文本通过insert_textbox插入,确保使用原样式
  • 插入文本时设置居中对齐,保证位置与原文本完全匹配
  • 保存文档时使用garbage=3,确保保留原文档的字体资源,避免字体缺失

内容的提问来源于stack exchange,提问作者Sik Saw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 02:02:05