使用Python+OpenCV无法提取特定PDF中二维码的解决方案求助
问题背景
我是Python和OpenCV的初学者,编写了一个简易应用的代码,实现先将PDF转换为PNG格式图像,再通过OpenCV提取二维码的功能。该代码对我测试的部分PDF文件可正常运行,但处理某一特定PDF文件时,完成格式转换后程序提示未找到二维码。
这是我目前遇到的唯一一个无法从中提取二维码的PDF文件,我的目标是找到可从该特定PDF文件中提取二维码并完成解码的解决方案。
以下是转换为图像后cv2识别二维码情况的截图:

我在Python中的实现代码如下:
import cv2 import os from wand.image import Image as wi PDFfile = wi(filename="Sources/certificatFaresBenslamaAr.pdf",resolution=400) Images = PDFfile.convert('png') ImageSequence = 1 for img in PDFfile.sequence: Image = wi(image = img) Image.save(filename="Image"+str(ImageSequence)+".png") ImageSequence += 1 # read the QRCODE image image = cv2.imread("image1.png") # initialize the cv2 QRCode detector qrCodeDetector = cv2.QRCodeDetector() # detect and decode decodedText, points, straight_qrcode = qrCodeDetector.detectAndDecode(image) #points is the output array of vertices of the found QR code quadrangle #straight qrcode # if there is a QR code if points is not None: # QR Code detected handling code print('Decoded data: ' + decodedText) nrOfPoints = len(points) print('Number of points: ' + str(nrOfPoints)) points = points[0] for i in range(len(points)): pt1 = [int(val) for val in points[i]] pt2 = [int(val) for val in points[(i + 1) % 4]] cv2.line(image, pt1, pt2, color=(255, 0, 0), thickness=3) print('Successfully saved') cv2.imshow('Detected QR code', image) cv2.imwrite('Generated/extractedQrcode.png', image) cv2.waitKey(0) cv2.destroyAllWindows() else: print("QR code not detected") # display the image with lines # length of bounding box
我的代码在其他PDF示例上运行效果良好,但目前仅需要解决上述特定PDF的二维码提取问题。以下是可被我的代码正常识别的PDF示例截图:
注:我注意到我正在扫描的这个PDF文件中包含可选中的二维码,相关截图如下:
解决方案
问题原因
你遇到的问题核心有三个:
- 代码存在大小写敏感问题:保存的文件名为
Image1.png,读取时用的是image1.png,Linux/macOS系统下会直接读取失败 - OpenCV自带的
QRCodeDetector识别率较低,对矢量二维码转栅格后产生的边缘毛刺、对比度偏差容忍度很低 - 该PDF中的二维码是矢量格式,转PNG时如果参数设置不当,会丢失清晰的边缘特征
修复步骤
1. 基础问题修复
首先统一文件读写的文件名大小写,避免读取错误。
2. 优化PDF转PNG参数
修改wand转图逻辑,开启抗锯齿,保证矢量二维码转栅格后边缘清晰:
PDFfile = wi(filename="Sources/certificatFaresBenslamaAr.pdf", resolution=400, anti_aliasing=True)
3. 更换高识别率的二维码检测库
替换OpenCV自带检测器为pyzbar,识别成功率提升明显。先安装依赖:pip install pyzbar
Linux系统需额外安装libzbar0:sudo apt install libzbar0
4. 增加图像预处理逻辑
读取图像后先做灰度化+二值化处理,过滤干扰信息,提升识别率。
完整修改后代码
import cv2 import os from wand.image import Image as wi from pyzbar.pyzbar import decode # 转换PDF PDFfile = wi(filename="Sources/certificatFaresBenslamaAr.pdf", resolution=400, anti_aliasing=True) ImageSequence = 1 save_path = "Image1.png" for img in PDFfile.sequence: Image = wi(image = img) Image.save(filename=f"Image{ImageSequence}.png") ImageSequence += 1 # 读取图像+预处理 image = cv2.imread(save_path) gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) # 二值化处理 _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) # 检测二维码 qrcodes = decode(thresh) if qrcodes: for qr in qrcodes: decodedText = qr.data.decode('utf-8') print(f"Decoded data: {decodedText}") # 画标注框 points = qr.polygon if len(points) == 4: pts = [list(pt) for pt in points] for i in range(4): cv2.line(image, pts[i], pts[(i+1)%4], color=(255,0,0), thickness=3) print('Successfully saved') cv2.imwrite('Generated/extractedQrcode.png', image) # 如需可视化可取消下方注释 # cv2.imshow('Detected QR code', image) # cv2.waitKey(0) # cv2.destroyAllWindows() else: print("QR code not detected")
最优方案(无需转图)
因为该PDF中的二维码是可选中的矢量格式,可直接用pymupdf库直接提取PDF中的二维码,不需要转栅格,识别准确率100%,运行效率更高。
内容的提问来源于stack exchange,提问作者Road to engineering
相关产品推荐
相关产品推荐

