PDF文本替换异常:修改后输出未更新及fitz链接识别问题咨询
PDF文本替换与链接修改问题
我尝试打开abc.pdf文件,将其中的“google.com”替换为用户输入的自定义链接。虽然文本替换操作显示成功,但生成的output.pdf文件内容仍与原文件一致,修改无法保存。另外,用fitz库修改链接时,发现只能识别带http://前缀的链接,没法直接替换“google.com”这类纯文本形式的内容。
以下是我尝试的三段Python代码:
第一段代码
import PyPDF2 import fitz from PyPDF2 import PdfReader, PdfWriter import requests # # Replace with the URL of the PDF you want to download # pdf_url = input("Enter the URL of the pdf file to download: ") # # # Replace with the link you want to replace the original links with new_link = input("Enter the link you want to replace the original links with: ") # # # Download the PDF file # response = requests.get(pdf_url) # with open("abc.pdf", "wb") as f: # f.write(response.content) with open('abc.pdf', 'rb') as file: reader = PyPDF2.PdfReader(file) writer = PyPDF2.PdfWriter() # pdf dosyasının tüm sayfalarını oku for page in range(len(reader.pages)): text = reader.pages[page].extract_text() print(text) # aranacak stringi bul if "google.com" in text: # stringi değiştir text = text.replace("google.com", new_link) # pdf dosyasını yeniden yaz print(text) writer.add_page(reader.pages[page]) with open('output.pdf', 'wb') as output: writer.write(output) file.close() output.close()
第二段代码
import fitz import requests # Replace with the URL of the PDF you want to download pdf_url = input("Enter the URL of the pdf file to download: ") # Replace with the link you want to replace the original links with new_link = input("Enter the link you want to replace the original links with: ") old_link = input("Enter the link you want to replace ") # Download the PDF file response = requests.get(pdf_url) with open("file.pdf", "wb") as f: f.write(response.content) # Open the PDF and modify the links pdf_doc = fitz.open("file.pdf") for page in pdf_doc: for link in page.links(): print(link) if "uri" in link and link["uri"] == old_link: print("Found one") link["uri"] = new_link # Save the modified PDF to the desktop pdf_doc.save("test2.pdf") pdf_doc.close()
第三段代码
import PyPDF2 import fitz from PyPDF2 import PdfReader, PdfWriter import requests # # Replace with the URL of the PDF you want to download # pdf_url = input("Enter the URL of the pdf file to download: ") # # # Replace with the link you want to replace the original links with new_link = input("Enter the link you want to replace the original links with: ") # # # Download the PDF file # response = requests.get(pdf_url) # with open("abc.pdf", "wb") as f: # f.write(response.content) # Open the original PDF file # with open('abc.pdf', 'rb') as file: doc = fitz.open('abc.pdf') print(doc) p = fitz.Point(50, 72) # start point of 1st line for page in doc: print(page) text = page.get_text() text = text.replace("google.com", new_link).encode("utf8") rc = page.insert_text(p, # bottom-left of 1st char text, # the text (honors '\n') fontname="helv", # the default font fontsize=11, # the default font size rotate=0, # also available: 90, 180, 270 ) # print(text) # page.set_text(text) # doc.insert_pdf(text,to_page=0) doc.save("output.pdf") doc.close()
问题分析与解决方法
第一段代码问题
你仅提取文本并修改了内存中的字符串,但未将修改后的文本写回PDF页面。PyPDF2的PdfReader提取的文本是只读的,修改字符串不会影响原页面内容,你直接把原页面添加到PdfWriter,所以输出文件和原文件一致。
第二段代码问题
page.links()只能识别PDF里的超链接对象(带URI属性),纯文本的“google.com”不属于链接对象,因此这段代码仅能替换带http://的可点击链接,无法处理纯文本内容。
第三段代码问题
insert_text是在页面指定位置插入新文本,而非替换原有文本,所以你只是在页面左上角插入了修改后的文本,原内容仍保留,且每次循环用同一坐标会导致文本重叠。
正确实现方案(针对纯文本“google.com”替换)
使用fitz(PyMuPDF)的文本定位与替换功能,可精准替换并保留原有格式:
import fitz new_link = input("输入要替换成的链接: ") doc = fitz.open("abc.pdf") for page in doc: # 查找所有包含"google.com"的文本区域 text_instances = page.search_for("google.com") for inst in text_instances: # 用白色覆盖原有文本,避免残留 page.draw_rect(inst, color=(1,1,1), fill=(1,1,1)) # 在原位置插入新文本 page.insert_text(fitz.Point(inst.x0, inst.y0), new_link, fontsize=11, color=(0,0,1)) doc.save("output.pdf") doc.close()
若需将替换后的文本转为可点击链接,可在插入文本后添加链接属性:
import fitz new_link = input("输入要替换成的链接: ") doc = fitz.open("abc.pdf") for page in doc: text_instances = page.search_for("google.com") for inst in text_instances: page.draw_rect(inst, color=(1,1,1), fill=(1,1,1)) # 插入新文本 page.insert_text(fitz.Point(inst.x0, inst.y0), new_link, fontsize=11, color=(0,0,1)) # 添加超链接,区域与原文本位置一致 page.insert_link({ "kind": fitz.LINK_URI, "uri": new_link, "rect": inst }) doc.save("output.pdf") doc.close()
内容的提问来源于stack exchange,提问作者hobik
相关产品推荐
相关产品推荐

