You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyPDF2移除PDF文本后转PNG出现空白问题咨询

问题:移除PDF文本后转PNG显示空白

我想要从带文本的PDF生成不含文本的PDF,使用了以下Python程序:

def remove_text_from_pdf(pdf_path_in, pdf_path_out):
    '''Removes the text from the PDF file and saves it as a new PDF file'''
    #Open the PDF file with the diagram and the text in read mode
    pdf_file = open(pdf_path_in, 'rb')
    
    #Create a PDF reader and writer object
    pdf_reader = PyPDF2.PdfReader(pdf_file) #OLDER VERSION WAS IN USE
    pdf_writer = PyPDF2.PdfWriter(pdf_file) #OLDER VERSION 'PDFFILEWRITER' IN USE

    #Get the pages from the PDF reader
    page = pdf_reader.pages[0]

    #Add the pages from the pdf reader to the pdf writer
    pdf_writer.add_page(page)

    #Remove the text from all pages added to the writer
    pdf_writer.remove_text()

    #Open the text output file in write mode
    out_file = open(pdf_path_out, 'wb')

    #Save the information to the text file
    pdf_writer.write(out_file)

    return

随后使用以下函数将输出文件转为PNG:

def convert_pdf_to_png(pdf_path, png_path):
    '''Converts a PDF file to a PNG file'''
    #Set the image maximum pixels to be none so that it doesn't give a DOS attack error
    pdffile = pdf_path
    doc = fitz.open(pdffile)
    page = doc.load_page(0)  # number of page
    pix = page.get_pixmap()
    output = png_path
    pix.save(output)
    doc.close()

但得到的PNG文件是纯空白的白色图像,预期生成的PDF应为非空白状态,请问该如何解决?


问题原因与解决方案

核心问题分析

  • PyPDF2初始化错误:创建PdfWriter时传入原PDF的读流pdf_file是错误用法,PdfWriter应初始化为空对象,绑定原文件流会导致生成的PDF结构损坏,最终转PNG显示空白。
  • remove_text()方法局限性:PyPDF2的该方法仅能移除标准文本层,若原PDF文本嵌入矢量图形、图像中,或是扫描件格式,方法无法生效甚至破坏页面渲染。
  • 文件流未正确关闭:代码未显式关闭文件流,可能导致文件写入不完整。

修复方案

方案1:修正PyPDF2代码

调整PdfWriter初始化方式,用with语句自动管理文件流避免资源泄漏:

import PyPDF2

def remove_text_from_pdf(pdf_path_in, pdf_path_out):
    '''Removes the text from the PDF file and saves it as a new PDF file'''
    with open(pdf_path_in, 'rb') as pdf_file:
        pdf_reader = PyPDF2.PdfReader(pdf_file)
        # 初始化空的PdfWriter,不再传入原文件流
        pdf_writer = PyPDF2.PdfWriter()

        page = pdf_reader.pages[0]
        pdf_writer.add_page(page)

        pdf_writer.remove_text()

        with open(pdf_path_out, 'wb') as out_file:
            pdf_writer.write(out_file)

方案2:使用PyMuPDF(fitz)直接处理(更可靠)

PyMuPDF对PDF结构操作更稳定,可直接移除页面文本块后保存PDF,再转PNG:

import fitz

def remove_text_from_pdf(pdf_path_in, pdf_path_out):
    '''使用PyMuPDF移除PDF文本'''
    doc = fitz.open(pdf_path_in)
    page = doc.load_page(0)
    
    # 获取所有文本块并用白色覆盖
    text_blocks = page.get_text("blocks")
    for block in text_blocks:
        rect = fitz.Rect(block[:4])
        page.add_redact_annot(rect, fill=(1,1,1))
    page.apply_redactions()
    
    doc.save(pdf_path_out)
    doc.close()

def convert_pdf_to_png(pdf_path, png_path):
    '''Converts a PDF file to a PNG file'''
    doc = fitz.open(pdf_path)
    page = doc.load_page(0)
    pix = page.get_pixmap(dpi=300)  # 指定DPI提升清晰度
    pix.save(png_path)
    doc.close()

验证步骤

  1. 运行修复后的remove_text_from_pdf生成新PDF
  2. 手动打开新PDF确认内容非空白且文本已移除
  3. 运行convert_pdf_to_png转PNG,检查输出图像是否正常

内容的提问来源于stack exchange,提问作者Rutvij Gholap

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 19:43:03