You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python批量清理PDF文件,仅保留指定目标图片?

解决方案:用PyMuPDF(fitz)批量处理PDF

PyPDF对PDF底层图形、图像的支持有限,而**PyMuPDF(fitz)**是处理这类需求的理想工具——它能直接访问PDF页面的所有底层对象(文本、路径、图像、形状等),完全满足你批量清理、保留指定图片的需求。

核心思路

因为你的PDF结构一致,可通过两种方式定位目标图片:

  • 基于固定位置/尺寸:利用结构一致的特点,直接匹配目标图的坐标范围或像素尺寸(最稳定)
  • 基于颜色特征:提取图像的颜色直方图,筛选黄色像素占比达标的图片(适合黄色标记直接在图片上的场景)

具体实现代码

以下是批量处理的示例代码,以“固定位置匹配”为例:

import fitz
import os

def clean_pdf(input_path, output_path, target_img_bbox):
    """
    清理PDF,仅保留指定区域内的图片
    :param input_path: 输入PDF路径
    :param output_path: 输出PDF路径
    :param target_img_bbox: 目标图片的边界框 (x0, y0, x1, y1),可通过先分析单个PDF获取
    """
    doc = fitz.open(input_path)
    for page in doc:
        # 清空当前页面所有原有内容
        page.clean_contents()
        # 获取页面所有图像对象
        img_list = page.get_images(full=True)
        for img_index in img_list:
            xref = img_index[0]
            # 获取图像的实际位置边界框
            img_rect = page.get_image_bbox(img_index)
            # 判断图像是否在目标区域内(可根据需求调整匹配精度)
            if fitz.Rect(target_img_bbox).contains(img_rect):
                # 将目标图像重新插入原位置
                page.insert_image(img_rect, xref=xref)
        # 彻底删除页面残留的文本和路径形状
        page.delete_text()
        page.delete_shapes()
    doc.save(output_path)
    doc.close()

# 批量处理文件夹下的所有PDF文件
def batch_clean_pdfs(input_dir, output_dir, target_img_bbox):
    if not os.path.exists(output_dir):
        os.makedirs(output_dir)
    for filename in os.listdir(input_dir):
        if filename.lower().endswith(".pdf"):
            input_path = os.path.join(input_dir, filename)
            output_path = os.path.join(output_dir, f"cleaned_{filename}")
            clean_pdf(input_path, output_path, target_img_bbox)
            print(f"处理完成:{filename}")

# ----------------------
# 先运行这段代码获取目标图的边界框(仅需执行一次)
# test_doc = fitz.open("sample.pdf")
# test_page = test_doc[0]
# for img in test_page.get_images(full=True):
#     print(test_page.get_image_bbox(img))  # 输出的坐标就是target_img_bbox
# test_doc.close()
# ----------------------

# 调用批量处理(替换为你的实际路径和目标坐标)
batch_clean_pdfs("input_pdfs", "output_cleaned", (100, 200, 500, 600))

针对黄色标记图片的优化

如果目标图片是通过黄色路径(比如黄色框)标记的,可以先遍历页面的路径元素,识别黄色(RGB约为(1,1,0)或相近值)的框,再匹配框内的图像:

def get_yellow_bboxes(page):
    yellow_bboxes = []
    for shape in page.get_drawings():
        # 检查路径的填充色或描边色是否为黄色
        if shape["fill"] == (1,1,0) or shape["color"] == (1,1,0):
            yellow_bboxes.append(fitz.Rect(shape["rect"]))
    return yellow_bboxes

随后在clean_pdf函数中,用这些黄色框来匹配图像位置即可。

注意事项

  • 若PDF存在加密,需先调用doc.authenticate("password")解密
  • 扫描版PDF(整页为单张图像)可直接按位置筛选,无需额外处理
  • 测试时先拿单个PDF验证逻辑,再批量运行

内容的提问来源于stack exchange,提问作者Kai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 14:01:03