You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu环境下Pytesseract图像文本识别不准确问题求助

Ubuntu 22.04下Pytesseract文本识别失真问题

我在Ubuntu 22.04机器上运行Python程序,流程为:用PyAutoGUI截取屏幕区域,再通过OpenCV和Pytesseract识别截图中的文本。这套程序在Windows 10上能完美识别,但在Ubuntu上的识别结果严重失真。

待识别图像:
待识别图像

所用Python代码

import cv2
from PIL import Image
import pytesseract

img = cv2.imread("screenshot.png")

text = pytesseract.image_to_string(img, lang="eng")
print(text)

注:Windows版本仅需配置Tesseract可执行文件路径,逻辑完全一致。

识别结果对比

  • Windows机器输出:
748 - Cry Me A River
  • Ubuntu机器输出:
VL} ea wa ow WN Aois

补充信息

  • 两台机器分辨率均为1920x1080
  • 按相同像素尺寸截图,识别结果仍不同
  • 使用Windows生成的截图在两台机器上识别,结果仍有差异
  • Windows机器硬件配置高于Ubuntu机器(不确定是否相关)

解决建议

  1. 检查Tesseract版本与语言包

    • 执行 tesseract --version 查看Ubuntu上的版本,尽量和Windows端版本保持一致(版本差异可能导致识别模型适配性问题)。
    • 确认已安装英文语言包:运行 sudo apt install tesseract-ocr-eng,避免语言包缺失或损坏。
  2. 图像预处理优化
    OpenCV默认读取BGR格式图像,而Pytesseract更适配RGB格式,同时增加二值化处理提升文本对比度:

    import cv2
    import pytesseract
    
    # 读取图像并转换为RGB格式
    img = cv2.imread("screenshot.png")
    img_rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
    
    # 灰度化+二值化预处理
    img_gray = cv2.cvtColor(img_rgb, cv2.COLOR_RGB2GRAY)
    _, img_thresh = cv2.threshold(img_gray, 127, 255, cv2.THRESH_BINARY)
    
    text = pytesseract.image_to_string(img_thresh, lang="eng")
    print(text)
    
  3. 调整截图与图像读取方式
    PyAutoGUI在Ubuntu上的截图可能存在色彩空间差异,尝试改用PIL直接保存和读取截图:

    import pyautogui
    from PIL import Image
    import pytesseract
    
    # 用PIL保存截图
    screenshot = pyautogui.screenshot()
    screenshot.save("screenshot.png")
    
    # 直接用PIL读取图像传入Pytesseract
    img = Image.open("screenshot.png")
    text = pytesseract.image_to_string(img, lang="eng")
    print(text)
    
  4. 指定Tesseract识别参数
    添加配置参数限定识别范围和模式,减少干扰:

    text = pytesseract.image_to_string(img_thresh, lang="eng", config="--psm 6 -c tessedit_char_whitelist=0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz- ")
    

    --psm 6 表示假设图像为单一均匀文本块,tessedit_char_whitelist 限定识别字符范围,过滤无关干扰项。


内容的提问来源于stack exchange,提问作者Alfred Langer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 08:12:18