Ubuntu环境下Pytesseract图像文本识别不准确问题求助
Ubuntu 22.04下Pytesseract文本识别失真问题
我在Ubuntu 22.04机器上运行Python程序,流程为:用PyAutoGUI截取屏幕区域,再通过OpenCV和Pytesseract识别截图中的文本。这套程序在Windows 10上能完美识别,但在Ubuntu上的识别结果严重失真。
待识别图像:
所用Python代码
import cv2 from PIL import Image import pytesseract img = cv2.imread("screenshot.png") text = pytesseract.image_to_string(img, lang="eng") print(text)
注:Windows版本仅需配置Tesseract可执行文件路径,逻辑完全一致。
识别结果对比
- Windows机器输出:
748 - Cry Me A River
- Ubuntu机器输出:
VL} ea wa ow WN Aois
补充信息
- 两台机器分辨率均为1920x1080
- 按相同像素尺寸截图,识别结果仍不同
- 使用Windows生成的截图在两台机器上识别,结果仍有差异
- Windows机器硬件配置高于Ubuntu机器(不确定是否相关)
解决建议
检查Tesseract版本与语言包
- 执行
tesseract --version查看Ubuntu上的版本,尽量和Windows端版本保持一致(版本差异可能导致识别模型适配性问题)。 - 确认已安装英文语言包:运行
sudo apt install tesseract-ocr-eng,避免语言包缺失或损坏。
- 执行
图像预处理优化
OpenCV默认读取BGR格式图像,而Pytesseract更适配RGB格式,同时增加二值化处理提升文本对比度:import cv2 import pytesseract # 读取图像并转换为RGB格式 img = cv2.imread("screenshot.png") img_rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) # 灰度化+二值化预处理 img_gray = cv2.cvtColor(img_rgb, cv2.COLOR_RGB2GRAY) _, img_thresh = cv2.threshold(img_gray, 127, 255, cv2.THRESH_BINARY) text = pytesseract.image_to_string(img_thresh, lang="eng") print(text)调整截图与图像读取方式
PyAutoGUI在Ubuntu上的截图可能存在色彩空间差异,尝试改用PIL直接保存和读取截图:import pyautogui from PIL import Image import pytesseract # 用PIL保存截图 screenshot = pyautogui.screenshot() screenshot.save("screenshot.png") # 直接用PIL读取图像传入Pytesseract img = Image.open("screenshot.png") text = pytesseract.image_to_string(img, lang="eng") print(text)指定Tesseract识别参数
添加配置参数限定识别范围和模式,减少干扰:text = pytesseract.image_to_string(img_thresh, lang="eng", config="--psm 6 -c tessedit_char_whitelist=0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz- ")--psm 6表示假设图像为单一均匀文本块,tessedit_char_whitelist限定识别字符范围,过滤无关干扰项。
内容的提问来源于stack exchange,提问作者Alfred Langer
相关产品推荐
相关产品推荐

