如何用Python批量清理PDF文件,仅保留指定目标图片?
解决方案:用PyMuPDF(fitz)批量处理PDF
PyPDF对PDF底层图形、图像的支持有限,而**PyMuPDF(fitz)**是处理这类需求的理想工具——它能直接访问PDF页面的所有底层对象(文本、路径、图像、形状等),完全满足你批量清理、保留指定图片的需求。
核心思路
因为你的PDF结构一致,可通过两种方式定位目标图片:
- 基于固定位置/尺寸:利用结构一致的特点,直接匹配目标图的坐标范围或像素尺寸(最稳定)
- 基于颜色特征:提取图像的颜色直方图,筛选黄色像素占比达标的图片(适合黄色标记直接在图片上的场景)
具体实现代码
以下是批量处理的示例代码,以“固定位置匹配”为例:
import fitz import os def clean_pdf(input_path, output_path, target_img_bbox): """ 清理PDF,仅保留指定区域内的图片 :param input_path: 输入PDF路径 :param output_path: 输出PDF路径 :param target_img_bbox: 目标图片的边界框 (x0, y0, x1, y1),可通过先分析单个PDF获取 """ doc = fitz.open(input_path) for page in doc: # 清空当前页面所有原有内容 page.clean_contents() # 获取页面所有图像对象 img_list = page.get_images(full=True) for img_index in img_list: xref = img_index[0] # 获取图像的实际位置边界框 img_rect = page.get_image_bbox(img_index) # 判断图像是否在目标区域内(可根据需求调整匹配精度) if fitz.Rect(target_img_bbox).contains(img_rect): # 将目标图像重新插入原位置 page.insert_image(img_rect, xref=xref) # 彻底删除页面残留的文本和路径形状 page.delete_text() page.delete_shapes() doc.save(output_path) doc.close() # 批量处理文件夹下的所有PDF文件 def batch_clean_pdfs(input_dir, output_dir, target_img_bbox): if not os.path.exists(output_dir): os.makedirs(output_dir) for filename in os.listdir(input_dir): if filename.lower().endswith(".pdf"): input_path = os.path.join(input_dir, filename) output_path = os.path.join(output_dir, f"cleaned_{filename}") clean_pdf(input_path, output_path, target_img_bbox) print(f"处理完成:{filename}") # ---------------------- # 先运行这段代码获取目标图的边界框(仅需执行一次) # test_doc = fitz.open("sample.pdf") # test_page = test_doc[0] # for img in test_page.get_images(full=True): # print(test_page.get_image_bbox(img)) # 输出的坐标就是target_img_bbox # test_doc.close() # ---------------------- # 调用批量处理(替换为你的实际路径和目标坐标) batch_clean_pdfs("input_pdfs", "output_cleaned", (100, 200, 500, 600))
针对黄色标记图片的优化
如果目标图片是通过黄色路径(比如黄色框)标记的,可以先遍历页面的路径元素,识别黄色(RGB约为(1,1,0)或相近值)的框,再匹配框内的图像:
def get_yellow_bboxes(page): yellow_bboxes = [] for shape in page.get_drawings(): # 检查路径的填充色或描边色是否为黄色 if shape["fill"] == (1,1,0) or shape["color"] == (1,1,0): yellow_bboxes.append(fitz.Rect(shape["rect"])) return yellow_bboxes
随后在clean_pdf函数中,用这些黄色框来匹配图像位置即可。
注意事项
- 若PDF存在加密,需先调用
doc.authenticate("password")解密 - 扫描版PDF(整页为单张图像)可直接按位置筛选,无需额外处理
- 测试时先拿单个PDF验证逻辑,再批量运行
内容的提问来源于stack exchange,提问作者Kai
相关产品推荐
相关产品推荐

