You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用开源工具替换PDF文本并保留原样式(含iText7问题)

使用PyMuPDF实现带样式的PDF文本替换

相比iText7,PyMuPDF(Python绑定)在提取文本样式和保留样式替换的操作上更简洁,能满足你替换指定文本并完全保留原字体、字号、颜色及装饰样式(加粗、斜体、下划线)的需求。

步骤1:安装PyMuPDF

pip install pymupdf

步骤2:实现替换代码

import fitz  # PyMuPDF

def replace_text_with_style(input_pdf, output_pdf, replacements):
    # 打开目标PDF
    doc = fitz.open(input_pdf)
    for page in doc:
        # 以字典格式提取页面所有文本及样式信息
        text_dict = page.get_text("dict")
        for block in text_dict["blocks"]:
            if "lines" not in block:
                continue
            # 遍历每一行文本
            for line in block["lines"]:
                # 遍历行内的每个文本片段(每个片段对应统一样式)
                for span in line["spans"]:
                    original_text = span["text"]
                    # 匹配需要替换的文本
                    if original_text in replacements:
                        new_text = replacements[original_text]
                        # 提取原文本的样式参数
                        font_name = span["font"]
                        font_size = span["size"]
                        text_color = span["color"]
                        style_flags = span["flags"]
                        text_bbox = span["bbox"]  # 文本位置边界框

                        # 删除原文本
                        page.add_redact_annot(text_bbox, text="")
                        page.apply_redactions()

                        # 解析样式标识:flags包含加粗(2)、斜体(1)
                        is_bold = (style_flags & 2) != 0
                        is_italic = (style_flags & 1) != 0
                        # 检查是否有下划线
                        has_underline = span.get("underline", False)

                        # 插入带原样式的新文本
                        text_writer = fitz.TextWriter(page.rect)
                        text_writer.append(
                            fitz.Point(text_bbox[0], text_bbox[1]),
                            new_text,
                            fontname=font_name,
                            fontsize=font_size,
                            color=fitz.sRGB_to_pdf(text_color),
                            bold=is_bold,
                            italic=is_italic
                        )
                        text_writer.write_text(page)

                        # 手动添加下划线(PyMuPDF暂不直接支持下划线样式继承)
                        if has_underline:
                            underline_y = text_bbox[3] + 1  # 下划线位置微调
                            page.draw_line(
                                fitz.Point(text_bbox[0], underline_y),
                                fitz.Point(text_bbox[2], underline_y),
                                color=fitz.sRGB_to_pdf(text_color),
                                width=0.5
                            )
    # 保存修改后的PDF
    doc.save(output_pdf)
    doc.close()

# 定义替换规则:原文本 -> 新文本
replace_rules = {
    "write": "check",
    "java": "perl",
    "texts": "words"
}

# 执行替换
replace_text_with_style("input.pdf", "output.pdf", replace_rules)

代码说明

  1. 文本样式提取:通过page.get_text("dict")获取每个文本片段的字体、字号、颜色、样式标识及位置,确保能完全继承原样式。
  2. 文本替换逻辑:先删除原文本,再插入应用了原样式的新文本,保证位置和样式一致。
  3. 下划线处理:由于下划线属于文本装饰属性,需手动绘制线条模拟,位置基于原文本边界框微调。

内容的提问来源于stack exchange,提问作者Erfan Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 02:23:24