You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pytesseract无法识别小尺寸裁剪图片中的数字,求解决方案

问题描述

有一张从原图裁剪得到的小尺寸数字图片,尝试图片放大、阈值处理等预处理操作后,仍无法通过Pytesseract识别其中的数字。

原识别代码:

import cv2
import pytesseract
from pytesseract import Output

img = cv2.imread('rois/roi11.jpg')
data = pytesseract.image_to_boxes(img, output_type=Output.DICT)
print(data)

尝试的预处理代码:

import cv2 
import pytesseract
img = cv2.imread('rois/roi11.jpg')
img2 = cv2.resize(img, (0, 0), fx=2, fy=2)
gry = cv2.cvtColor(img2, cv2.COLOR_BGR2GRAY)
thr = cv2.threshold(gry, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)[1]
data = pytesseract.image_to_string(thr)
print(data)
可行解决方法
  • 指定数字识别模式:Pytesseract默认识别全字符,通过--psm 10(单字符识别模式)和-c tessedit_char_whitelist=0123456789(限定仅识别数字)缩小识别范围,提升精准度。同时改用INTER_CUBIC插值放大,保留更多字符细节:
import cv2
import pytesseract

img = cv2.imread('rois/roi11.jpg')
# 用三次插值放大4倍,比线性插值更清晰
img_enlarged = cv2.resize(img, None, fx=4, fy=4, interpolation=cv2.INTER_CUBIC)
gray = cv2.cvtColor(img_enlarged, cv2.COLOR_BGR2GRAY)
# 自适应阈值处理,适配局部明暗差异
thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
# 设置自定义识别参数
custom_config = r'--oem 3 --psm 10 -c tessedit_char_whitelist=0123456789'
result = pytesseract.image_to_string(thresh, config=custom_config)
print(result.strip())
  • 形态学优化轮廓:对阈值处理后的图像进行膨胀操作,填补数字边缘的细小缺口,强化字符轮廓:
import cv2
import pytesseract
import numpy as np

img = cv2.imread('rois/roi11.jpg')
img_enlarged = cv2.resize(img, None, fx=4, fy=4, interpolation=cv2.INTER_CUBIC)
gray = cv2.cvtColor(img_enlarged, cv2.COLOR_BGR2GRAY)
thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
# 创建2x2结构元素,执行膨胀操作
kernel = np.ones((2, 2), np.uint8)
dilated_img = cv2.dilate(thresh, kernel, iterations=1)
custom_config = r'--oem 3 --psm 10 -c tessedit_char_whitelist=0123456789'
result = pytesseract.image_to_string(dilated_img, config=custom_config)
print(result.strip())
  • 更换OCR引擎模式:若你的Tesseract版本支持,设置--oem 1启用LSTM神经网络引擎,对模糊、小字符的识别效果优于传统引擎:
custom_config = r'--oem 1 --psm 10 -c tessedit_char_whitelist=0123456789'
  • 降噪预处理:如果图片存在明显噪点,先通过高斯模糊降噪,再进行后续处理:
gray = cv2.cvtColor(img_enlarged, cv2.COLOR_BGR2GRAY)
# 高斯模糊降噪
blurred = cv2.GaussianBlur(gray, (3, 3), 0)
thresh = cv2.adaptiveThreshold(blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)

内容的提问来源于stack exchange,提问作者naiveprogrammer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 17:55:20