You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升基于pytesseract的图像文本提取准确率?

图像文本提取优化方案

核心问题排查

  • 你当前代码中生成了对比度增强后的图像img_contrasted,但实际识别时仍使用原图像im,这是对比度优化无效的直接原因。
  • 选用的--psm 12模式(稀疏文本识别)不匹配当前规整的文本布局,这类布局更适合针对结构化文本的PSM模式。

具体优化步骤

1. 修正图像使用逻辑

将识别对象替换为增强后的图像,确保对比度优化生效:

# 替换原识别代码中的im为img_contrasted
extracted_text = pytesseract.image_to_string(img_contrasted, lang='eng', config=custom_config)
data = pytesseract.image_to_data(img_contrasted, output_type=Output.DICT, lang='eng', config=custom_config)
data1 = pytesseract.image_to_data(img_contrasted, lang='eng', config=custom_config)

2. 调整Tesseract PSM模式

推荐使用以下两种模式之一,适配当前图像的规整文本结构:

  • --psm 6:假设图像是单一统一的文本块(最适合当前场景)
  • --psm 3:默认自动模式,适合多数通用文本场景

修改配置代码:

# 改用PSM 6模式
custom_config = r'--oem 3 --psm 6'

3. 增加二值化与去噪处理

针对印刷体文本,二值化(转黑白)能进一步提升识别率,可在对比度增强后添加以下步骤:

# 二值化处理(阈值可根据实际调整)
threshold = 127
img_binary = img_contrasted.point(lambda p: p > threshold and 255)
# 后续识别使用img_binary
extracted_text = pytesseract.image_to_string(img_binary, lang='eng', config=custom_config)

4. 补充字符白名单(可选)

如果图像中仅包含英文字母、空格和标点,可添加字符白名单减少识别干扰:

custom_config = r'--oem 3 --psm 6 --tessedit_char_whitelist ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz '

完整优化后代码示例

import pytesseract
from PIL import Image, ImageEnhance
from pytesseract import Output

custom_config = r'--oem 3 --psm 6 --tessedit_char_whitelist ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz '
pytesseract.pytesseract.tesseract_cmd = r"C:\Users\LENOVO\AppData\Local\Programs\Tesseract-OCR\tesseract.exe"

# 加载图像
im = Image.open("C:\\Users\\LENOVO\\Desktop\\photo2.png")
im = im.convert("RGB")

# 对比度增强
curr_con = ImageEnhance.Contrast(im)
img_contrasted = curr_con.enhance(4.0)

# 二值化处理
threshold = 127
img_binary = img_contrasted.point(lambda p: p > threshold and 255)

# 执行识别
extracted_text = pytesseract.image_to_string(img_binary, lang='eng', config=custom_config)
data = pytesseract.image_to_data(img_binary, output_type=Output.DICT, lang='eng', config=custom_config)

print("提取文本:")
print(extracted_text)
print("\n详细识别数据:")
print(data)

额外建议

  • 确保你的Tesseract-OCR安装包包含完整的英文语言包,若缺失可重新下载安装最新版本。
  • 如果图像存在倾斜,可添加图像旋转校正步骤(使用PIL的Image.rotate)。

内容的提问来源于stack exchange,提问作者prisha sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 20:35:04