You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中通过预处理ROI提升Pytesseract的OCR识别精度

优化ROI预处理以提升Pytesseract文本识别精度

我正在开发一个项目,需要对屏幕上的感兴趣区域(ROI)进行预处理,以此提升Pytesseract的文本识别精度。目前已经实现了屏幕捕获以及基于OpenCV模板匹配的ROI定位功能,但ROI的预处理效果不佳,导致Pytesseract出现识别错误——比如会把数字「3,415」识别成「bais」。

核心处理函数

以下是用于预处理图像并输出pytesseract.image_to_string()结果的核心代码:

### images是字典,frame是pyautogui.screenshot()捕获的帧对应的numpy数组 ###
def draw_boxes(images, frame):
    rectangle_color = (0, 255, 0)

    ### 遍历images中的模板图,在屏幕帧中搜索匹配 ###
    for name, img in images.items():

        gray_frame = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
        gray_img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)

        result = cv2.matchTemplate(gray_frame, gray_img, cv2.TM_CCOEFF_NORMED)
        min_val, max_val, min_loc, max_loc = cv2.minMaxLoc(result)

        ### 如果模板匹配成功,定义感兴趣区域(ROI) ###
        if max_val > 0.5:
            ### 在CV2窗口绘制矩形框用于调试 ###
            top_left = max_loc
            bottom_right = (top_left[0] + img.shape[1], top_left[1] + img.shape[0])
            cv2.rectangle(frame, top_left, bottom_right, rectangle_color, 2)

            ### 定义感兴趣区域(ROI) ###
            x, y, w, h = max_loc[0], max_loc[1], img.shape[1], img.shape[0]
            roi = frame[y:y+h, x:x+w]

            ### 预处理感兴趣区域 ###
            roi_gray = cv2.cvtColor(roi, cv2.COLOR_BGR2GRAY)
            blurred_image = cv2.medianBlur(roi_gray, 1, 5)
            denoised_image = cv2.fastNlMeansDenoising(roi_gray, None, 12, 7, 21)
            roi_processed = denoised_image

            ### 应用二值化阈值,供Pytesseract识别 ###
            # inverted_image = cv2.bitwise_not(roi_processed)
            otsu_threshold_value, thresh_image = cv2.threshold(roi_processed, 1, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)

            ### 显示OCR视图用于调试开发 ###
            cv2.imshow('OCR View', thresh_image)
            cv2.waitKey(1)

            custom_config = r'--oem 3 --psm 11'
            text = pytesseract.image_to_string(thresh_image, lang='eng', config=custom_config)
            
    try:
        print(text)
    except UnboundLocalError:
        print('unbound')

Pytesseract识别输出示例

控制台输出中可见明显识别错误:

Buy Offer

Ectoplasm

Gx

:

A substance from those passing through the Underworld.

2

2

3379

Buy limit per 4 hours: 25,000

Quantity

Your price per item:

25,000

**bais**

85,375,000 coins

You have bought 1 so far for 3,415 coins.

125,000

可以看到Pytesseract将目标数字「3,415」误识别为「bais」。我认为需要优化预处理流程,让处理后的图像更清晰,但不清楚具体调整方向。

预处理前后的ROI对比

预处理前的ROI

预处理前的ROI

预处理后的ROI

预处理后的ROI

内容的提问来源于stack exchange,提问作者XENMS ACC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 01:07:49