如何检测PDF产品标签页中缺失的产品图与条形码图片?
解决PDF页面图片缺失的准确检测方案
核心问题梳理
你的每页PDF包含14个产品标签,每个标签对应2张图(产品图+条形码),因此每页预期有效图片总数为28张。之前仅通过奇偶性判断的方法,在缺失偶数张图时会完全失效,且原脚本的统计逻辑可能未覆盖所有真实存在的图片。
改进方案
1. 修正统计逻辑,精准匹配预期数量
原脚本仅依赖page.images且只过滤宽高>0的图片,可能遗漏部分嵌入图片。可以结合page.objects['image']获取更全面的图片数据,同时去重+过滤无效占位图,最后直接对比预期的28张数量:
import pdfplumber pdf_file = 'file.pdf' EXPECTED_IMAGES_PER_PAGE = 14 * 2 # 14个标签×2张图/标签 with pdfplumber.open(pdf_file) as pdf: for page_num, page in enumerate(pdf.pages, start=1): # 合并两种渠道的图片数据,避免遗漏 all_images = page.images + page.objects.get('image', []) # 去重+过滤有效图片(用坐标+ID标识避免重复统计,设置最小宽高过滤占位图) valid_images = [] seen_img_keys = set() for img in all_images: img_key = (img.get('x0'), img.get('top'), img.get('width'), img.get('height'), img.get('object_id')) if img_key not in seen_img_keys and img.get('width', 0) > 10 and img.get('height', 0) > 10: seen_img_keys.add(img_key) valid_images.append(img) actual_count = len(valid_images) # 直接对比预期数量,替代奇偶判断 if actual_count != EXPECTED_IMAGES_PER_PAGE: print(f"第{page_num}页图片异常:实际{actual_count}张,预期{EXPECTED_IMAGES_PER_PAGE}张")
2. 精准定位缺失标签(进阶方案)
如果需要定位到具体哪个标签缺图,可以利用PDF的固定布局(8行2列),预先计算每个标签的坐标区域,检查每个区域内的图片数量是否为2:
import pdfplumber pdf_file = 'file.pdf' # 先获取页面基础尺寸(假设所有页尺寸一致) with pdfplumber.open(pdf_file) as pdf: first_page = pdf.pages[0] page_w, page_h = first_page.width, first_page.height # 配置布局参数(根据你的PDF实际边距调整) ROW_NUM = 8 COL_NUM = 2 MARGIN_TOP = 50 MARGIN_BOTTOM = 50 MARGIN_LEFT = 30 MARGIN_RIGHT = 30 # 计算单个标签的宽高 tag_w = (page_w - MARGIN_LEFT - MARGIN_RIGHT) / COL_NUM tag_h = (page_h - MARGIN_TOP - MARGIN_BOTTOM) / ROW_NUM with pdfplumber.open(pdf_file) as pdf: for page_num, page in enumerate(pdf.pages, start=1): print(f"=== 第{page_num}页标签检测 ===") for row in range(ROW_NUM): for col in range(COL_NUM): # 计算当前标签的坐标范围(预留5px容错空间) x0 = MARGIN_LEFT + col * tag_w top = MARGIN_TOP + row * tag_h x1 = x0 + tag_w bottom = top + tag_h # 筛选该区域内的有效图片 tag_images = [ img for img in page.images if img['x0'] >= x0 - 5 and img['x1'] <= x1 + 5 and img['top'] >= top - 5 and img['bottom'] <= bottom + 5 and img['width'] > 10 and img['height'] > 10 ] if len(tag_images) != 2: print(f"第{row+1}行第{col+1}列标签异常:实际{len(tag_images)}张,预期2张")
额外提示
- 代码中的边距、最小宽高阈值(10px)需要根据你的PDF实际情况调整
- 如果pdfplumber仍无法捕获所有图片,可以尝试
PyMuPDF(fitz)库,它对PDF图片的提取兼容性更强
内容的提问来源于stack exchange,提问作者Jerome Cremades
相关产品推荐
相关产品推荐

