You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+OpenCV识别PDF中QR码时遇UnicodeDecodeError问题求助

从发票PDF的QR码提取数据时的Unicode解码错误解决

问题描述

我编写了一个类,用于从带有QR码的发票PDF中提取IBAN等债权人数据,原本运行正常,但处理某份PDF时出现如下错误:

UnicodeDecodeError: 'utf-8' codec can't decode byte 0xfc in position 157: invalid start byte

实现逻辑:使用fitz打开PDF,将每页渲染为PNG图片,再用OpenCV的QRCodeDetector识别QR码内容,从中提取IBAN和参考编号。核心代码如下:

doc = fitz.open(self.image_path)  # open document
i = 0

if not os.path.exists(f"./qr_codes/"):
    os.makedirs(f"./qr_codes/")

for page in doc:
    pix = page.get_pixmap(matrix=self.mat)  # render page to an image
    pix.save(f"./qr_codes/page_{i}.png")

    img = cv2.imread(f"./qr_codes/page_{i}.png")
    detect = cv2.QRCodeDetector()
    text, points, straight_qrcode = detect.detectAndDecode(img)

    if text:
        # 提取IBAN
        self.iban = "\r\n".join([line for line in text.splitlines() if re.findall(r"CH\d{19}", line.strip())])
        # 提取参考编号(仅保留数字用于SAP)
        ref_number = re.findall(r'CH\s*QRR\s*\d+|$', " ".join(text.splitlines()))
        self.ref_number = int(re.sub(r"\D","", ref_number[0])) if ref_number else None
        self.__save_values()
        return True
    i += 1
return False

尝试用numpy数组处理图片但仅得到空文本,代码如下:

stream = open(f'./qr_codes/page_{i}.png', encoding="utf-8", errors="ignore")
    stream = bytearray(stream.read(), encoding="utf-8")
    detect = cv2.QRCodeDetector()
    text, points, straight_qrcode = detect.detectAndDecode(numpy.asarray(stream, dtype=numpy.uint8))
    # print(text)

完整报错回溯:

Traceback (most recent call last):
  File "C:\Users\m7073\Repos\Chronos_New\invoice_extraction\qr_code_scan.py", line 128, in <module>
    qrcode.set_qr_values()
  File "C:\Users\m7073\Repos\Chronos_New\invoice_extraction\qr_code_scan.py", line 73, in set_qr_values
    text, points, straight_qrcode = detect.detectAndDecode(img)
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xfc in position 157: invalid start byte

最小复现代码:

import cv2
img = cv2.imread(f"page_1.png")
detect = cv2.QRCodeDetector()
text, points, straight_qrcode = detect.detectAndDecode(img)

解决方法

1. 使用detectAndDecodeMulti获取原始字节并指定编码解码

错误中的0xfc对应拉丁1(ISO-8859-1)编码的ü字符,说明QR码内容并非UTF-8编码。改用detectAndDecodeMulti方法可获取原始字节数据,再手动指定编码解码:

import cv2

img = cv2.imread("page_1.png")
detect = cv2.QRCodeDetector()
# 检测并获取原始解码信息
retval, decoded_info, _, _ = detect.detectAndDecodeMulti(img)

if retval:
    for info in decoded_info:
        # 若为字节类型,用ISO-8859-1解码(适配瑞士发票常见特殊字符)
        if isinstance(info, bytes):
            text = info.decode('ISO-8859-1')
        else:
            text = info
        # 后续提取IBAN和参考编号的逻辑
        print(text)

2. 强制解码并跳过错误字符

如果不确定编码格式,可在解码时设置errors参数跳过或替换错误字节,避免程序崩溃:

import cv2

img = cv2.imread("page_1.png")
detect = cv2.QRCodeDetector()
_, _, straight_qrcode = detect.detectAndDecode(img)

if straight_qrcode is not None:
    # 从straight_qrcode提取原始字节
    raw_bytes = straight_qrcode.tobytes()
    # 用replace替换无法解码的字符
    text = raw_bytes.decode('utf-8', errors='replace')
    # 后续处理逻辑
    print(text)

3. 修正numpy数组处理图片的方式

PNG是二进制文件,不能用文本编码打开,正确的处理方式如下:

import cv2
import numpy as np

# 二进制模式读取图片文件
with open('./qr_codes/page_1.png', 'rb') as f:
    img_data = f.read()
# 转换为numpy数组并解码为图片
nparr = np.frombuffer(img_data, np.uint8)
img = cv2.imdecode(nparr, cv2.IMREAD_COLOR)

detect = cv2.QRCodeDetector()
text, _, _ = detect.detectAndDecode(img)
# 尝试编码转换
if text:
    try:
        # 尝试将latin1编码转为utf-8
        text = text.encode('latin1').decode('utf-8')
    except UnicodeError:
        # 失败则强制替换错误字符
        text = text.encode('utf-8', errors='replace').decode('utf-8')
    print(text)

4. 优化PDF渲染参数

调整fitz的渲染参数,确保QR码清晰无信息丢失:

# 使用2倍分辨率矩阵渲染,提高清晰度
mat = fitz.Matrix(2.0, 2.0)
# 指定RGB颜色空间,避免灰度模式导致的识别问题
pix = page.get_pixmap(matrix=mat, colorspace=fitz.csRGB)
pix.save(f"./qr_codes/page_{i}.png")

内容的提问来源于stack exchange,提问作者user3793935

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 12:55:03