You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python OCR中‘data type = 19 is not supported’错误?

解决Python OCR中"data type = 19 is not supported"错误

嘿,这个错误我之前也碰到过!问题出在你定义形态学操作核的代码行——np.ones((1,1), np.unicode)使用了错误的数据类型,OpenCV的膨胀(dilation)和腐蚀(erosion)操作不支持np.unicode类型的核,它需要的是数值型数组(比如np.uint8)。

错误原因解析

data type = 19对应的正是np.unicode的类型编码,OpenCV无法处理这种类型的数组作为形态学操作的输入,因此抛出了这个不支持的错误。

修正后的完整代码

import pytesseract
from PIL import Image, ImageEnhance, ImageFilter
import cv2
import numpy as np

src_path = ""  # 这里填你的图片路径

def get_string(img_path):
    # 用OpenCV读取图片
    img = cv2.imread(img_path)
    # 转灰度图
    img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    
    # 修正:用np.uint8类型创建形态学核
    kernel = np.ones((1, 1), np.uint8)
    # 膨胀操作(补全你没写完的代码)
    img = cv2.dilate(img, kernel, iterations=1)
    # 腐蚀操作去噪
    img = cv2.erode(img, kernel, iterations=1)
    
    # 把OpenCV格式的图片转成PIL格式,方便pytesseract处理
    img_pil = Image.fromarray(img)
    
    # 可选:增强图像对比度,提升OCR识别率
    enhancer = ImageEnhance.Contrast(img_pil)
    img_pil = enhancer.enhance(2)
    
    # 用pytesseract提取文本
    text = pytesseract.image_to_string(img_pil, lang='chi_sim')  # 中文识别加lang参数,英文可省略
    return text

# 测试调用
if __name__ == "__main__":
    result = get_string(src_path)
    print("提取到的文本:")
    print(result)

额外优化建议

  • 如果噪点比较多,可以适当调大核的尺寸,比如np.ones((2,2), np.uint8),增强去噪效果
  • 如果图片模糊,可以先加个滤波操作,比如img = cv2.GaussianBlur(img, (3,3), 0)
  • 确保已经安装了Tesseract的语言包(比如中文识别需要chi_sim包)

内容的提问来源于stack exchange,提问作者harmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:27:32