You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+OpenCV无法提取特定PDF中二维码的解决方案求助

问题背景

我是Python和OpenCV的初学者,编写了一个简易应用的代码,实现先将PDF转换为PNG格式图像,再通过OpenCV提取二维码的功能。该代码对我测试的部分PDF文件可正常运行,但处理某一特定PDF文件时,完成格式转换后程序提示未找到二维码。
这是我目前遇到的唯一一个无法从中提取二维码的PDF文件,我的目标是找到可从该特定PDF文件中提取二维码并完成解码的解决方案。
以下是转换为图像后cv2识别二维码情况的截图:
转换后二维码识别情况截图
转换后二维码识别情况截图
我在Python中的实现代码如下:

import cv2
import os
from wand.image import Image as wi
PDFfile = wi(filename="Sources/certificatFaresBenslamaAr.pdf",resolution=400)
Images = PDFfile.convert('png')
ImageSequence = 1

for img in PDFfile.sequence:
    Image = wi(image = img)
    Image.save(filename="Image"+str(ImageSequence)+".png")
    ImageSequence += 1

    # read the QRCODE image
image = cv2.imread("image1.png")
# initialize the cv2 QRCode detector
qrCodeDetector = cv2.QRCodeDetector()
# detect and decode
decodedText, points, straight_qrcode = qrCodeDetector.detectAndDecode(image)
#points is the output array of vertices of the found QR code quadrangle
#straight qrcode
# if there is a QR code
if points is not None:
    # QR Code detected handling code

    print('Decoded data: ' + decodedText)
    nrOfPoints = len(points)
    print('Number of points:  ' + str(nrOfPoints))
    points = points[0]
    for i in range(len(points)):
        pt1 = [int(val) for val in points[i]]
        pt2 = [int(val) for val in points[(i + 1) % 4]]
        cv2.line(image, pt1, pt2, color=(255, 0, 0), thickness=3)
  
    print('Successfully saved')
    cv2.imshow('Detected QR code', image)
    cv2.imwrite('Generated/extractedQrcode.png', image)
    cv2.waitKey(0)
    cv2.destroyAllWindows()
else:
    print("QR code not detected")
    # display the image with lines
    # length of bounding box

我的代码在其他PDF示例上运行效果良好,但目前仅需要解决上述特定PDF的二维码提取问题。以下是可被我的代码正常识别的PDF示例截图:
可正常识别的PDF示例截图
注:我注意到我正在扫描的这个PDF文件中包含可选中的二维码,相关截图如下:
可选中二维码的PDF截图

解决方案

问题原因

你遇到的问题核心有三个:

  1. 代码存在大小写敏感问题:保存的文件名为Image1.png,读取时用的是image1.png,Linux/macOS系统下会直接读取失败
  2. OpenCV自带的QRCodeDetector识别率较低,对矢量二维码转栅格后产生的边缘毛刺、对比度偏差容忍度很低
  3. 该PDF中的二维码是矢量格式,转PNG时如果参数设置不当,会丢失清晰的边缘特征

修复步骤

1. 基础问题修复

首先统一文件读写的文件名大小写,避免读取错误。

2. 优化PDF转PNG参数

修改wand转图逻辑,开启抗锯齿,保证矢量二维码转栅格后边缘清晰:

PDFfile = wi(filename="Sources/certificatFaresBenslamaAr.pdf", resolution=400, anti_aliasing=True)

3. 更换高识别率的二维码检测库

替换OpenCV自带检测器为pyzbar,识别成功率提升明显。先安装依赖:
pip install pyzbar
Linux系统需额外安装libzbar0:sudo apt install libzbar0

4. 增加图像预处理逻辑

读取图像后先做灰度化+二值化处理,过滤干扰信息,提升识别率。

完整修改后代码

import cv2
import os
from wand.image import Image as wi
from pyzbar.pyzbar import decode

# 转换PDF
PDFfile = wi(filename="Sources/certificatFaresBenslamaAr.pdf", resolution=400, anti_aliasing=True)
ImageSequence = 1
save_path = "Image1.png"
for img in PDFfile.sequence:
    Image = wi(image = img)
    Image.save(filename=f"Image{ImageSequence}.png")
    ImageSequence += 1

# 读取图像+预处理
image = cv2.imread(save_path)
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
# 二值化处理
_, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)

# 检测二维码
qrcodes = decode(thresh)
if qrcodes:
    for qr in qrcodes:
        decodedText = qr.data.decode('utf-8')
        print(f"Decoded data: {decodedText}")
        # 画标注框
        points = qr.polygon
        if len(points) == 4:
            pts = [list(pt) for pt in points]
            for i in range(4):
                cv2.line(image, pts[i], pts[(i+1)%4], color=(255,0,0), thickness=3)
    print('Successfully saved')
    cv2.imwrite('Generated/extractedQrcode.png', image)
    # 如需可视化可取消下方注释
    # cv2.imshow('Detected QR code', image)
    # cv2.waitKey(0)
    # cv2.destroyAllWindows()
else:
    print("QR code not detected")

最优方案(无需转图)

因为该PDF中的二维码是可选中的矢量格式,可直接用pymupdf库直接提取PDF中的二维码,不需要转栅格,识别准确率100%,运行效率更高。


内容的提问来源于stack exchange,提问作者Road to engineering

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 06:48:03