You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测PDF产品标签页中缺失的产品图与条形码图片?

解决PDF页面图片缺失的准确检测方案

核心问题梳理

你的每页PDF包含14个产品标签,每个标签对应2张图(产品图+条形码),因此每页预期有效图片总数为28张。之前仅通过奇偶性判断的方法,在缺失偶数张图时会完全失效,且原脚本的统计逻辑可能未覆盖所有真实存在的图片。

改进方案

1. 修正统计逻辑,精准匹配预期数量

原脚本仅依赖page.images且只过滤宽高>0的图片,可能遗漏部分嵌入图片。可以结合page.objects['image']获取更全面的图片数据,同时去重+过滤无效占位图,最后直接对比预期的28张数量:

import pdfplumber

pdf_file = 'file.pdf'
EXPECTED_IMAGES_PER_PAGE = 14 * 2  # 14个标签×2张图/标签

with pdfplumber.open(pdf_file) as pdf:
    for page_num, page in enumerate(pdf.pages, start=1):
        # 合并两种渠道的图片数据,避免遗漏
        all_images = page.images + page.objects.get('image', [])
        # 去重+过滤有效图片(用坐标+ID标识避免重复统计,设置最小宽高过滤占位图)
        valid_images = []
        seen_img_keys = set()
        for img in all_images:
            img_key = (img.get('x0'), img.get('top'), img.get('width'), img.get('height'), img.get('object_id'))
            if img_key not in seen_img_keys and img.get('width', 0) > 10 and img.get('height', 0) > 10:
                seen_img_keys.add(img_key)
                valid_images.append(img)
        actual_count = len(valid_images)
        
        # 直接对比预期数量,替代奇偶判断
        if actual_count != EXPECTED_IMAGES_PER_PAGE:
            print(f"第{page_num}页图片异常:实际{actual_count}张,预期{EXPECTED_IMAGES_PER_PAGE}张")

2. 精准定位缺失标签(进阶方案)

如果需要定位到具体哪个标签缺图,可以利用PDF的固定布局(8行2列),预先计算每个标签的坐标区域,检查每个区域内的图片数量是否为2:

import pdfplumber

pdf_file = 'file.pdf'
# 先获取页面基础尺寸(假设所有页尺寸一致)
with pdfplumber.open(pdf_file) as pdf:
    first_page = pdf.pages[0]
    page_w, page_h = first_page.width, first_page.height

# 配置布局参数(根据你的PDF实际边距调整)
ROW_NUM = 8
COL_NUM = 2
MARGIN_TOP = 50
MARGIN_BOTTOM = 50
MARGIN_LEFT = 30
MARGIN_RIGHT = 30

# 计算单个标签的宽高
tag_w = (page_w - MARGIN_LEFT - MARGIN_RIGHT) / COL_NUM
tag_h = (page_h - MARGIN_TOP - MARGIN_BOTTOM) / ROW_NUM

with pdfplumber.open(pdf_file) as pdf:
    for page_num, page in enumerate(pdf.pages, start=1):
        print(f"=== 第{page_num}页标签检测 ===")
        for row in range(ROW_NUM):
            for col in range(COL_NUM):
                # 计算当前标签的坐标范围(预留5px容错空间)
                x0 = MARGIN_LEFT + col * tag_w
                top = MARGIN_TOP + row * tag_h
                x1 = x0 + tag_w
                bottom = top + tag_h
                # 筛选该区域内的有效图片
                tag_images = [
                    img for img in page.images
                    if img['x0'] >= x0 - 5 and img['x1'] <= x1 + 5
                    and img['top'] >= top - 5 and img['bottom'] <= bottom + 5
                    and img['width'] > 10 and img['height'] > 10
                ]
                if len(tag_images) != 2:
                    print(f"第{row+1}行第{col+1}列标签异常:实际{len(tag_images)}张,预期2张")

额外提示

  • 代码中的边距、最小宽高阈值(10px)需要根据你的PDF实际情况调整
  • 如果pdfplumber仍无法捕获所有图片,可以尝试PyMuPDF(fitz)库,它对PDF图片的提取兼容性更强

内容的提问来源于stack exchange,提问作者Jerome Cremades

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 03:32:38