You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTesseract无法识别LCD屏幕文字,多种预处理尝试无效求助

问题:PyTesseract能检测LCD区域但无法识别文字

我在图像预处理阶段尝试了多种操作,但PyTesseract仍无法识别LCD屏幕上的文字。它能在LCD周围生成bounding box,说明检测到了区域,但无法输出文字内容。

原始图像:
原始图像

我的代码如下:

import cv2
import pytesseract
import numpy as np

img = cv2.imread("test-python2.jpg")

gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)

ret, thresh1 = cv2.threshold(gray, 50, 255, cv2.THRESH_OTSU | cv2.THRESH_BINARY_INV)

rect_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (18, 18))
kernel = np.ones((5, 5), np.uint8)
#closin = cv2.morphologyEx(gray, cv2.MORPH_CLOSE, kernel)

dilation = cv2.dilate(thresh1, rect_kernel, iterations = 1)

contours, hierarchy = cv2.findContours(dilation, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)
im2 = img.copy()
for cnt in contours:
    x, y, w, h = cv2.boundingRect(cnt)
    im2 = cv2.rectangle(im2, (x, y), (x + w, y + h), (0, 255, 0), 2)
    
    cropped = img[y:y + h, x:x + w]
    text = pytesseract.image_to_string(cropped)
    im2 = cv2.putText(im2, text, (x, y - 10), cv2.FONT_HERSHEY_SIMPLEX, 1.2, (0, 255, 0), 3)
    text2 = text.encode('latin-1', 'replace').decode('latin-1')
    print (text2)


cv2.imshow("", im2)
cv2.waitKey(0)

cv2.imshow()的输出图像:
处理后输出图像

目前其他区域的文字识别结果足够准确,但就是无法识别LCD屏幕上的内容。我尝试过多种二值化和阈值处理方法,但始终无法成功识别LCD文字,而LCD识别对我的项目至关重要,我已在此问题上卡壳许久,恳请帮助。


解决方案

针对LCD屏幕文字识别的问题,核心是LCD文字的对比度干扰、底色差异以及字符特征与常规印刷体的区别导致识别失效,可通过以下步骤优化:

1. 针对性预处理LCD区域

当前全局预处理无法适配LCD的蓝绿色底色,需单独对裁剪后的LCD区域做精细化处理:

# 替换原代码中cropped后的识别逻辑
cropped_gray = cv2.cvtColor(cropped, cv2.COLOR_BGR2GRAY)
# 自适应阈值处理,适配LCD局部亮度波动
adaptive_thresh = cv2.adaptiveThreshold(cropped_gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
# 中值滤波去除噪点
adaptive_thresh = cv2.medianBlur(adaptive_thresh, 3)

2. 给PyTesseract添加专属识别参数

LCD文字多为等宽数字/简单符号,通过限定字符集和识别模式可大幅提升准确率:

# --psm 7 指定单一行文本模式;白名单限定识别数字和小数点
text = pytesseract.image_to_string(adaptive_thresh, config='--psm 7 -c tessedit_char_whitelist=0123456789.')

3. 微调形态学操作强化字符边缘

LCD字符较纤细,用小核做轻微膨胀可强化字符轮廓:

kernel_small = np.ones((2,2), np.uint8)
adaptive_thresh = cv2.dilate(adaptive_thresh, kernel_small, iterations=1)

完整优化后的核心代码片段

for cnt in contours:
    x, y, w, h = cv2.boundingRect(cnt)
    im2 = cv2.rectangle(im2, (x, y), (x + w, y + h), (0, 255, 0), 2)
    
    cropped = img[y:y + h, x:x + w]
    # LCD专属预处理
    cropped_gray = cv2.cvtColor(cropped, cv2.COLOR_BGR2GRAY)
    adaptive_thresh = cv2.adaptiveThreshold(cropped_gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
    adaptive_thresh = cv2.medianBlur(adaptive_thresh, 3)
    kernel_small = np.ones((2,2), np.uint8)
    adaptive_thresh = cv2.dilate(adaptive_thresh, kernel_small, iterations=1)
    # 带参数的OCR识别
    text = pytesseract.image_to_string(adaptive_thresh, config='--psm 7 -c tessedit_char_whitelist=0123456789.')
    im2 = cv2.putText(im2, text, (x, y - 10), cv2.FONT_HERSHEY_SIMPLEX, 1.2, (0, 255, 0), 3)
    text2 = text.encode('latin-1', 'replace').decode('latin-1')
    print(text2)

额外建议

如果上述方法仍未达到预期,可尝试:

  • 对LCD区域做透视变换,校正角度偏移;
  • 针对你的LCD字体训练Tesseract自定义字符集。

内容的提问来源于stack exchange,提问作者skullx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 02:15:58