You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF文本替换异常:修改后输出未更新及fitz链接识别问题咨询

PDF文本替换与链接修改问题

我尝试打开abc.pdf文件,将其中的“google.com”替换为用户输入的自定义链接。虽然文本替换操作显示成功,但生成的output.pdf文件内容仍与原文件一致,修改无法保存。另外,用fitz库修改链接时,发现只能识别带http://前缀的链接,没法直接替换“google.com”这类纯文本形式的内容。

以下是我尝试的三段Python代码:

第一段代码

import PyPDF2
import fitz
from PyPDF2 import PdfReader, PdfWriter
import requests

# # Replace with the URL of the PDF you want to download
# pdf_url = input("Enter the URL of the pdf file to download: ")
#
# # Replace with the link you want to replace the original links with
new_link = input("Enter the link you want to replace the original links with: ")
#
# # Download the PDF file
# response = requests.get(pdf_url)
# with open("abc.pdf", "wb") as f:
#     f.write(response.content)
with open('abc.pdf', 'rb') as file:
    reader = PyPDF2.PdfReader(file)
    writer = PyPDF2.PdfWriter()

    # pdf dosyasının tüm sayfalarını oku
    for page in range(len(reader.pages)):
        text = reader.pages[page].extract_text()
        print(text)
        # aranacak stringi bul
        if "google.com" in text:
            # stringi değiştir
            text = text.replace("google.com", new_link)
            # pdf dosyasını yeniden yaz
            print(text)
            writer.add_page(reader.pages[page])
            with open('output.pdf', 'wb') as output:
                writer.write(output)
file.close()
output.close()

第二段代码

import fitz
import requests

# Replace with the URL of the PDF you want to download
pdf_url = input("Enter the URL of the pdf file to download: ")

# Replace with the link you want to replace the original links with
new_link = input("Enter the link you want to replace the original links with: ")
old_link = input("Enter the link you want to replace  ")

# Download the PDF file
response = requests.get(pdf_url)
with open("file.pdf", "wb") as f:
    f.write(response.content)

# Open the PDF and modify the links
pdf_doc = fitz.open("file.pdf")
for page in pdf_doc:
    for link in page.links():
        print(link)
        if "uri" in link and link["uri"] == old_link:
            print("Found one")
            link["uri"] = new_link

# Save the modified PDF to the desktop
pdf_doc.save("test2.pdf")
pdf_doc.close()

第三段代码

import PyPDF2
import fitz
from PyPDF2 import PdfReader, PdfWriter
import requests

# # Replace with the URL of the PDF you want to download
# pdf_url = input("Enter the URL of the pdf file to download: ")
#
# # Replace with the link you want to replace the original links with
new_link = input("Enter the link you want to replace the original links with: ")
#
# # Download the PDF file
# response = requests.get(pdf_url)
# with open("abc.pdf", "wb") as f:
#     f.write(response.content)
# Open the original PDF file
# with open('abc.pdf', 'rb') as file:
doc = fitz.open('abc.pdf')
print(doc)
p = fitz.Point(50, 72)  # start point of 1st line

for page in doc:
    print(page)
    text = page.get_text()
    text = text.replace("google.com", new_link).encode("utf8")
    rc = page.insert_text(p,  # bottom-left of 1st char
                          text,  # the text (honors '\n')
                          fontname="helv",  # the default font
                          fontsize=11,  # the default font size
                          rotate=0,  # also available: 90, 180, 270
                          )    # print(text)
    # page.set_text(text)
    # doc.insert_pdf(text,to_page=0)
doc.save("output.pdf")
doc.close()

问题分析与解决方法

第一段代码问题

你仅提取文本并修改了内存中的字符串,但未将修改后的文本写回PDF页面。PyPDF2的PdfReader提取的文本是只读的,修改字符串不会影响原页面内容,你直接把原页面添加到PdfWriter,所以输出文件和原文件一致。

第二段代码问题

page.links()只能识别PDF里的超链接对象(带URI属性),纯文本的“google.com”不属于链接对象,因此这段代码仅能替换带http://的可点击链接,无法处理纯文本内容。

第三段代码问题

insert_text是在页面指定位置插入新文本,而非替换原有文本,所以你只是在页面左上角插入了修改后的文本,原内容仍保留,且每次循环用同一坐标会导致文本重叠。


正确实现方案(针对纯文本“google.com”替换)

使用fitz(PyMuPDF)的文本定位与替换功能,可精准替换并保留原有格式:

import fitz

new_link = input("输入要替换成的链接: ")
doc = fitz.open("abc.pdf")

for page in doc:
    # 查找所有包含"google.com"的文本区域
    text_instances = page.search_for("google.com")
    for inst in text_instances:
        # 用白色覆盖原有文本,避免残留
        page.draw_rect(inst, color=(1,1,1), fill=(1,1,1))
        # 在原位置插入新文本
        page.insert_text(fitz.Point(inst.x0, inst.y0), new_link, fontsize=11, color=(0,0,1))

doc.save("output.pdf")
doc.close()

若需将替换后的文本转为可点击链接,可在插入文本后添加链接属性:

import fitz

new_link = input("输入要替换成的链接: ")
doc = fitz.open("abc.pdf")

for page in doc:
    text_instances = page.search_for("google.com")
    for inst in text_instances:
        page.draw_rect(inst, color=(1,1,1), fill=(1,1,1))
        # 插入新文本
        page.insert_text(fitz.Point(inst.x0, inst.y0), new_link, fontsize=11, color=(0,0,1))
        # 添加超链接,区域与原文本位置一致
        page.insert_link({
            "kind": fitz.LINK_URI,
            "uri": new_link,
            "rect": inst
        })

doc.save("output.pdf")
doc.close()

内容的提问来源于stack exchange,提问作者hobik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 19:40:32