为何Tesseract 4.1.1在清晰数字时钟图像上识别性能不佳?
Tesseract 4.1.1数字时钟时间识别返回空结果的问题排查
我正在使用Tesseract 4.1.1开展数字时钟时间识别项目,但多数情况下无法识别内容,返回空结果。即便图像数据相当清晰,识别性能仍很差,明明Tesseract在更复杂任务上表现尚可,我已耗费大量时间却陷入瓶颈,恳请提供解决思路!
已尝试的优化措施
- 通过参数
-c tessedit_char_whitelist=0123456789:.设置感兴趣字符白名单; - 将图像转换为黑字白底的二值图像;
- 使用约120张图像、500轮迭代微调英文模型,增加迭代次数也未改善效果;
- 根据tessdoc建议,将图像缩放至字符x-height处于20-30像素区间。
实现代码
import pytesseract import math import cv2 import numpy as np def read_time(filename): # 白名单字符设置 tesseract_config = '-c tessedit_char_whitelist=0123456789:. --psm 6' # 图像转二值图 clock_image = cv2.imread(filename) gray_image = cv2.cvtColor(clock_image, cv2.COLOR_BGR2GRAY) _, binary_image = cv2.threshold(gray_image,200,255,cv2.THRESH_BINARY) # 判断背景是否为白色,否则反转图像 is_white_background = False if not is_white_background: binary_image = cv2.bitwise_not(binary_image) # 自定义模型配置 custom_model_config = '-c tessedit_char_whitelist=0123456789:. -l my_custom_model --psm 6' # 缩放图像至字符高度20-30像素区间 cnt_black_pixels_per_row = np.sum(binary_image == 0, axis=1) non_zero_lines = [n for n,i in enumerate(cnt_black_pixels_per_row) if i > 0] # 获取字符上下边界估算当前高度 top, bottom = non_zero_lines[0], non_zero_lines[-1] current_height = abs(top - bottom) height_expected = 32 new_height = math.floor(binary_image.shape[0] * height_expected / current_height) new_shape = (binary_image.shape[1], new_height) final_image = cv2.resize(binary_image, new_shape, interpolation= cv2.INTER_LINEAR) # 四种方式尝试识别 clock_text_binary = pytesseract.image_to_string(binary_image, config=tesseract_config) clock_text_binary_custom_config = pytesseract.image_to_string(binary_image, config=custom_model_config) clock_text_resized = pytesseract.image_to_string(final_image, config=custom_model_config) clock_text_resized_custom_config = pytesseract.image_to_string(final_image, config=tesseract_config) print(clock_text_binary, clock_text_binary_custom_config, clock_text_resized, clock_text_resized_custom_config)
识别失败的示例图像




内容的提问来源于stack exchange,提问作者h1t5uj1
相关产品推荐
相关产品推荐

