You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tesseract OCR Python代码输出乱码或无内容问题求助

Tesseract OCR识别无输出/乱码问题排查

我正在学习使用Tesseract OCR包读取图像并识别图像中的字符,但运行代码时,要么没有输出内容,要么显示图像中不存在的随机乱码。请帮忙排查可能的原因。

相关代码

import cv2
import pytesseract
from pytesseract import Output
import re

# 假设image已通过cv2.imread加载
h, w, c = image.shape
boxes = pytesseract.image_to_boxes(image)
for i in boxes.splitlines():
    i = i.split(' ')
    image = cv2.rectangle(image, (int(i[1]), h-int(i[2])),
                        (int(i[3]), h-int(i[4])), (0,255,0), 2)

dict = pytesseract.image_to_data(image, output_type=Output.DICT)
keys = list(dict.keys())
boxes = len(dict['text'])
date = '^(0[1-9]| [12][0-9]|3[01])/(0[1-9]|1[012])/(19|20)\d\d$'
for i in range(boxes):
    if int(float(dict['conf'][i])) > 60:
        (x,y,w,h) = (dict['left'][i], dict['top'][i], dict['width'][i], dict['height'][i])
        image = cv2.rectangle(image, (x,y), (x+w,y+h), (0,255,0), 2)
        if re.match(date, dict['text'][i]):
            (x,y,w,h) = (dict['left'][i], dict['top'][i], dict['width'][i], dict['height'][i])
            image = cv2.rectangle(image, (x,y), (x+w,y+h), (0,255,0), 2)

config1 = r'--oem 3 --psm 8'

print(pytesseract.image_to_string(image, config=config1))
cv2.imshow('image', image)
cv2.waitKey(0)
cv2.destroyAllWindows()

可能的原因及解决方法

  • 未做图像预处理:Tesseract对输入图像质量要求较高,直接用彩色原图易识别失败。建议添加预处理步骤:

    1. 转灰度图:gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
    2. 二值化增强对比度:_, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
    3. 降噪(可选):blur = cv2.GaussianBlur(gray, (3,3), 0)
      用预处理后的图像执行识别操作,而非原始彩色图。
  • 识别用图像被修改:代码先在原始图像上画矩形框,再用带框图像执行识别,框线会干扰识别逻辑。应保留原始图像用于识别,创建副本画框:

    draw_image = image.copy()
    # 所有cv2.rectangle操作都作用于draw_image,而非原始image
    
  • PSM模式选择错误:--psm 8仅适用于单个单词的识别场景,若图像包含多段文本或日期,该模式会导致识别异常。建议调整:

    • 识别日期可尝试--psm 6(假设输入是单一均匀文本块)
    • 通用场景用默认的--psm 3(自动分页/分块)
  • 正则表达式错误:日期匹配正则中[12][0-9]前多了一个空格,无法匹配正常格式日期,修正为:

    date = r'^(0[1-9]|[12][0-9]|3[01])/(0[1-9]|1[012])/(19|20)\d\d$'
    
  • 语言包缺失:若图像含非英文文本(如中文),未安装对应语言包会导致乱码。安装对应语言包后,识别时指定语言参数:

    print(pytesseract.image_to_string(preprocessed_image, config=config1, lang='chi_sim'))
    
  • 置信度判断逻辑冗余:代码对同一文本框重复画框,无实际意义,可删除重复画框的代码。

内容的提问来源于stack exchange,提问作者miths17

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 00:37:34