Python+OpenCV识别PDF中QR码时遇UnicodeDecodeError问题求助
从发票PDF的QR码提取数据时的Unicode解码错误解决
问题描述
我编写了一个类,用于从带有QR码的发票PDF中提取IBAN等债权人数据,原本运行正常,但处理某份PDF时出现如下错误:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xfc in position 157: invalid start byte
实现逻辑:使用fitz打开PDF,将每页渲染为PNG图片,再用OpenCV的QRCodeDetector识别QR码内容,从中提取IBAN和参考编号。核心代码如下:
doc = fitz.open(self.image_path) # open document i = 0 if not os.path.exists(f"./qr_codes/"): os.makedirs(f"./qr_codes/") for page in doc: pix = page.get_pixmap(matrix=self.mat) # render page to an image pix.save(f"./qr_codes/page_{i}.png") img = cv2.imread(f"./qr_codes/page_{i}.png") detect = cv2.QRCodeDetector() text, points, straight_qrcode = detect.detectAndDecode(img) if text: # 提取IBAN self.iban = "\r\n".join([line for line in text.splitlines() if re.findall(r"CH\d{19}", line.strip())]) # 提取参考编号(仅保留数字用于SAP) ref_number = re.findall(r'CH\s*QRR\s*\d+|$', " ".join(text.splitlines())) self.ref_number = int(re.sub(r"\D","", ref_number[0])) if ref_number else None self.__save_values() return True i += 1 return False
尝试用numpy数组处理图片但仅得到空文本,代码如下:
stream = open(f'./qr_codes/page_{i}.png', encoding="utf-8", errors="ignore") stream = bytearray(stream.read(), encoding="utf-8") detect = cv2.QRCodeDetector() text, points, straight_qrcode = detect.detectAndDecode(numpy.asarray(stream, dtype=numpy.uint8)) # print(text)
完整报错回溯:
Traceback (most recent call last): File "C:\Users\m7073\Repos\Chronos_New\invoice_extraction\qr_code_scan.py", line 128, in <module> qrcode.set_qr_values() File "C:\Users\m7073\Repos\Chronos_New\invoice_extraction\qr_code_scan.py", line 73, in set_qr_values text, points, straight_qrcode = detect.detectAndDecode(img) UnicodeDecodeError: 'utf-8' codec can't decode byte 0xfc in position 157: invalid start byte
最小复现代码:
import cv2 img = cv2.imread(f"page_1.png") detect = cv2.QRCodeDetector() text, points, straight_qrcode = detect.detectAndDecode(img)
解决方法
1. 使用detectAndDecodeMulti获取原始字节并指定编码解码
错误中的0xfc对应拉丁1(ISO-8859-1)编码的ü字符,说明QR码内容并非UTF-8编码。改用detectAndDecodeMulti方法可获取原始字节数据,再手动指定编码解码:
import cv2 img = cv2.imread("page_1.png") detect = cv2.QRCodeDetector() # 检测并获取原始解码信息 retval, decoded_info, _, _ = detect.detectAndDecodeMulti(img) if retval: for info in decoded_info: # 若为字节类型,用ISO-8859-1解码(适配瑞士发票常见特殊字符) if isinstance(info, bytes): text = info.decode('ISO-8859-1') else: text = info # 后续提取IBAN和参考编号的逻辑 print(text)
2. 强制解码并跳过错误字符
如果不确定编码格式,可在解码时设置errors参数跳过或替换错误字节,避免程序崩溃:
import cv2 img = cv2.imread("page_1.png") detect = cv2.QRCodeDetector() _, _, straight_qrcode = detect.detectAndDecode(img) if straight_qrcode is not None: # 从straight_qrcode提取原始字节 raw_bytes = straight_qrcode.tobytes() # 用replace替换无法解码的字符 text = raw_bytes.decode('utf-8', errors='replace') # 后续处理逻辑 print(text)
3. 修正numpy数组处理图片的方式
PNG是二进制文件,不能用文本编码打开,正确的处理方式如下:
import cv2 import numpy as np # 二进制模式读取图片文件 with open('./qr_codes/page_1.png', 'rb') as f: img_data = f.read() # 转换为numpy数组并解码为图片 nparr = np.frombuffer(img_data, np.uint8) img = cv2.imdecode(nparr, cv2.IMREAD_COLOR) detect = cv2.QRCodeDetector() text, _, _ = detect.detectAndDecode(img) # 尝试编码转换 if text: try: # 尝试将latin1编码转为utf-8 text = text.encode('latin1').decode('utf-8') except UnicodeError: # 失败则强制替换错误字符 text = text.encode('utf-8', errors='replace').decode('utf-8') print(text)
4. 优化PDF渲染参数
调整fitz的渲染参数,确保QR码清晰无信息丢失:
# 使用2倍分辨率矩阵渲染,提高清晰度 mat = fitz.Matrix(2.0, 2.0) # 指定RGB颜色空间,避免灰度模式导致的识别问题 pix = page.get_pixmap(matrix=mat, colorspace=fitz.csRGB) pix.save(f"./qr_codes/page_{i}.png")
内容的提问来源于stack exchange,提问作者user3793935
相关产品推荐
相关产品推荐

