You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+PyMuPDF实现PDF替换文本与原文本的格式样式匹配?

PyMuPDF替换PDF文本保留原样式的实现方法

我用Python的PyMuPDF库做PDF文本搜索替换,功能能正常运行,但替换后的文本没法匹配原文本的颜色、字号等样式,求解决办法。当前使用的代码如下:

import os
import fitz

# Prompt user for input of file name
file_name_input = input("Enter a start or text of the file name: ")

# Get a list of PDF files in the current directory matching the file name input
pdf_files = [f for f in os.listdir() if f.lower().endswith('.pdf') and file_name_input.lower() in f.lower()]

if not pdf_files:
    print("No PDF files found matching the file name input")
else:
    # Prompt user for input of search and replace text
    search_replace_list = []
    while True:
        search_text = input("Enter the search text (leave blank to exit): ")
        if not search_text:
            break
        replace_text = input("Enter the replace text: ")
        search_replace_list.append((search_text, replace_text))

    for file_name in pdf_files:
        pdf_file = fitz.open(file_name)
        found = False
        for page in pdf_file:
            for search_text, replace_text in search_replace_list:
                draft = page.search_for(search_text.strip(), hit_max=16, quads=True, quads_tol=0.01)
                if draft:
                    found = True
                    for rect in draft:
                        annot = page.add_redact_annot(rect, text=replace_text)
                    page.apply_redactions()
                    page.apply_redactions(images=fitz.PDF_REDACT_IMAGE_NONE)

        if found:
            output_file_name = file_name[:-4] + '_modified.pdf'
            pdf_file.save(output_file_name, garbage=False, deflate=True, encryption=False)
            print(f"Changes saved to {output_file_name}")
        else:
            print(f"No search text found in {file_name}")

        pdf_file.close()

问题原因

当前代码使用add_redact_annot进行替换时,默认采用PyMuPDF的内置文本样式,没有继承原文本的字体、颜色、字号等属性,导致替换后的样式和原文本不一致。

解决思路

要保留原样式,需要先获取待替换文本的原始样式信息(字体名称、字号、颜色等),再在替换时应用这些样式:

  • 遍历页面文本块,匹配搜索文本并提取对应样式属性
  • 清除原文本所在区域内容
  • 使用原始样式插入替换文本

修改后的代码

import os
import fitz

# 匹配文本并获取其样式信息
def find_text_with_style(page, search_text):
    text_dict = page.get_text("dict")
    matches = []
    for block in text_dict["blocks"]:
        if "lines" not in block:
            continue
        for line in block["lines"]:
            for span in line["spans"]:
                if search_text in span["text"]:
                    # 计算匹配文本的位置矩形(系数0.6为字体宽度估算值,可按需调整)
                    char_width = span["size"] * 0.6
                    start_x = span["bbox"][0] + span["text"].index(search_text) * char_width
                    rect = fitz.Rect(
                        start_x,
                        span["bbox"][1],
                        start_x + len(search_text) * char_width,
                        span["bbox"][3]
                    )
                    matches.append({
                        "rect": rect,
                        "font": span["font"],
                        "size": span["size"],
                        "color": span["color"],
                        "flags": span["flags"]
                    })
    return matches

# 主逻辑
file_name_input = input("Enter a start or text of the file name: ")
pdf_files = [f for f in os.listdir() if f.lower().endswith('.pdf') and file_name_input.lower() in f.lower()]

if not pdf_files:
    print("No PDF files found matching the file name input")
else:
    search_replace_list = []
    while True:
        search_text = input("Enter the search text (leave blank to exit): ")
        if not search_text:
            break
        replace_text = input("Enter the replace text: ")
        search_replace_list.append((search_text, replace_text))

    for file_name in pdf_files:
        pdf_file = fitz.open(file_name)
        found = False
        for page in pdf_file:
            for search_text, replace_text in search_replace_list:
                matches = find_text_with_style(page, search_text)
                if matches:
                    found = True
                    # 先清除原文本区域
                    for match in matches:
                        page.add_redact_annot(match["rect"])
                    page.apply_redactions()
                    # 插入带原样式的替换文本
                    for match in matches:
                        # 转换颜色格式为PyMuPDF可用格式
                        color = fitz.sRGB_to_pdf(match["color"])
                        page.insert_text(
                            match["rect"].tl,  # 以原文本左上角为插入起点
                            replace_text,
                            fontname=match["font"],
                            fontsize=match["size"],
                            color=color,
                            flags=match["flags"]
                        )

        if found:
            output_file_name = file_name[:-4] + '_modified.pdf'
            pdf_file.save(output_file_name, garbage=False, deflate=True, encryption=False)
            print(f"Changes saved to {output_file_name}")
        else:
            print(f"No search text found in {file_name}")

        pdf_file.close()

注意事项

  • 文本位置计算:使用字号×0.6估算字符宽度,不同字体可能存在偏差,可根据实际情况调整系数
  • 字体兼容性:若PDF使用嵌入字体,PyMuPDF需能识别该字体才能正确应用样式,否则会 fallback 到默认字体
  • 多行文本匹配:当前代码仅支持单行内的文本匹配,若搜索文本跨多行,需额外处理文本块的拼接逻辑

内容的提问来源于stack exchange,提问作者Hetul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 22:02:21