如何移除PDF中的图片?批量PDF数字签名图片移除遇阻求助
移除PDF中绿色圈注数字签名图片的可行方法
你尝试用PyPDF2复制页面的方案无效,处理后出现乱码和文字丢失,问题出在PyPDF2对部分复杂PDF的兼容性不足。以下是两种可靠的解决方法:
方案1:用PyMuPDF精准移除目标图片
PyMuPDF(又名fitz)对PDF结构的解析能力更强,能避免乱码问题,还可通过特征筛选移除绿色圈注图片。
- 先安装依赖:
pip install pymupdf
- 执行以下代码(可根据圈注的实际特征调整筛选逻辑):
import fitz def remove_green_signature_images(input_pdf, output_pdf): doc = fitz.open(input_pdf) for page in doc: # 获取页面所有图片的完整信息 image_list = page.get_images(full=True) for img in image_list: xref = img[0] # 获取图片在页面上的位置矩形 img_rects = page.get_image_rects(xref) # 假设绿色圈注是小尺寸图形,这里用尺寸筛选,可替换为颜色特征判断 for rect in img_rects: if rect.width < 50 and rect.height < 50: page.delete_image(xref) break doc.save(output_pdf) doc.close() # 替换为你的文件路径 input_path = 'C:\\Users\\Usuario\\Downloads\\JG_1_01221-2020-0-1801-JR-LA-06.pdf' output_path = 'C:\\Users\\Usuario\\Desktop\\DEP\\Lats_cleaned.pdf' remove_green_signature_images(input_path, output_path)
如果需要靠颜色精准筛选,可以提取图片像素分析绿色占比,比如:
# 提取图片像素判断颜色的示例片段 pix = fitz.Pixmap(doc, xref) # 检查图片是否以绿色为主(RGB值中G通道占比高) green_pixel_count = 0 total_pixels = pix.width * pix.height for y in range(pix.height): for x in range(pix.width): r, g, b = pix.pixel(x, y) if g > r and g > b and g > 200: # 绿色分量明显高于其他通道 green_pixel_count += 1 if green_pixel_count / total_pixels > 0.5: page.delete_image(xref)
方案2:用白色矩形覆盖圈注区域
如果图片识别困难,直接用白色矩形覆盖绿色圈注的位置,效果等同于移除:
import fitz def cover_green_signatures(input_pdf, output_pdf): doc = fitz.open(input_pdf) for page in doc: image_list = page.get_images(full=True) for img in image_list: xref = img[0] img_rects = page.get_image_rects(xref) for rect in img_rects: if rect.width < 50 and rect.height < 50: # 用白色填充覆盖该区域 page.draw_rect(rect, color=(1, 1, 1), fill=(1, 1, 1), overlay=True) doc.save(output_pdf) doc.close() cover_green_signatures(input_path, output_path)
内容的提问来源于stack exchange,提问作者Felipe
相关产品推荐
相关产品推荐

