使用PyPDF2移除PDF文本后转PNG出现空白问题咨询
问题:移除PDF文本后转PNG显示空白
我想要从带文本的PDF生成不含文本的PDF,使用了以下Python程序:
def remove_text_from_pdf(pdf_path_in, pdf_path_out): '''Removes the text from the PDF file and saves it as a new PDF file''' #Open the PDF file with the diagram and the text in read mode pdf_file = open(pdf_path_in, 'rb') #Create a PDF reader and writer object pdf_reader = PyPDF2.PdfReader(pdf_file) #OLDER VERSION WAS IN USE pdf_writer = PyPDF2.PdfWriter(pdf_file) #OLDER VERSION 'PDFFILEWRITER' IN USE #Get the pages from the PDF reader page = pdf_reader.pages[0] #Add the pages from the pdf reader to the pdf writer pdf_writer.add_page(page) #Remove the text from all pages added to the writer pdf_writer.remove_text() #Open the text output file in write mode out_file = open(pdf_path_out, 'wb') #Save the information to the text file pdf_writer.write(out_file) return
随后使用以下函数将输出文件转为PNG:
def convert_pdf_to_png(pdf_path, png_path): '''Converts a PDF file to a PNG file''' #Set the image maximum pixels to be none so that it doesn't give a DOS attack error pdffile = pdf_path doc = fitz.open(pdffile) page = doc.load_page(0) # number of page pix = page.get_pixmap() output = png_path pix.save(output) doc.close()
但得到的PNG文件是纯空白的白色图像,预期生成的PDF应为非空白状态,请问该如何解决?
问题原因与解决方案
核心问题分析
- PyPDF2初始化错误:创建
PdfWriter时传入原PDF的读流pdf_file是错误用法,PdfWriter应初始化为空对象,绑定原文件流会导致生成的PDF结构损坏,最终转PNG显示空白。 remove_text()方法局限性:PyPDF2的该方法仅能移除标准文本层,若原PDF文本嵌入矢量图形、图像中,或是扫描件格式,方法无法生效甚至破坏页面渲染。- 文件流未正确关闭:代码未显式关闭文件流,可能导致文件写入不完整。
修复方案
方案1:修正PyPDF2代码
调整PdfWriter初始化方式,用with语句自动管理文件流避免资源泄漏:
import PyPDF2 def remove_text_from_pdf(pdf_path_in, pdf_path_out): '''Removes the text from the PDF file and saves it as a new PDF file''' with open(pdf_path_in, 'rb') as pdf_file: pdf_reader = PyPDF2.PdfReader(pdf_file) # 初始化空的PdfWriter,不再传入原文件流 pdf_writer = PyPDF2.PdfWriter() page = pdf_reader.pages[0] pdf_writer.add_page(page) pdf_writer.remove_text() with open(pdf_path_out, 'wb') as out_file: pdf_writer.write(out_file)
方案2:使用PyMuPDF(fitz)直接处理(更可靠)
PyMuPDF对PDF结构操作更稳定,可直接移除页面文本块后保存PDF,再转PNG:
import fitz def remove_text_from_pdf(pdf_path_in, pdf_path_out): '''使用PyMuPDF移除PDF文本''' doc = fitz.open(pdf_path_in) page = doc.load_page(0) # 获取所有文本块并用白色覆盖 text_blocks = page.get_text("blocks") for block in text_blocks: rect = fitz.Rect(block[:4]) page.add_redact_annot(rect, fill=(1,1,1)) page.apply_redactions() doc.save(pdf_path_out) doc.close() def convert_pdf_to_png(pdf_path, png_path): '''Converts a PDF file to a PNG file''' doc = fitz.open(pdf_path) page = doc.load_page(0) pix = page.get_pixmap(dpi=300) # 指定DPI提升清晰度 pix.save(png_path) doc.close()
验证步骤
- 运行修复后的
remove_text_from_pdf生成新PDF - 手动打开新PDF确认内容非空白且文本已移除
- 运行
convert_pdf_to_png转PNG,检查输出图像是否正常
内容的提问来源于stack exchange,提问作者Rutvij Gholap
相关产品推荐
相关产品推荐

