如何解决Python OCR中‘data type = 19 is not supported’错误?
解决Python OCR中"data type = 19 is not supported"错误
嘿,这个错误我之前也碰到过!问题出在你定义形态学操作核的代码行——np.ones((1,1), np.unicode)使用了错误的数据类型,OpenCV的膨胀(dilation)和腐蚀(erosion)操作不支持np.unicode类型的核,它需要的是数值型数组(比如np.uint8)。
错误原因解析
data type = 19对应的正是np.unicode的类型编码,OpenCV无法处理这种类型的数组作为形态学操作的输入,因此抛出了这个不支持的错误。
修正后的完整代码
import pytesseract from PIL import Image, ImageEnhance, ImageFilter import cv2 import numpy as np src_path = "" # 这里填你的图片路径 def get_string(img_path): # 用OpenCV读取图片 img = cv2.imread(img_path) # 转灰度图 img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 修正:用np.uint8类型创建形态学核 kernel = np.ones((1, 1), np.uint8) # 膨胀操作(补全你没写完的代码) img = cv2.dilate(img, kernel, iterations=1) # 腐蚀操作去噪 img = cv2.erode(img, kernel, iterations=1) # 把OpenCV格式的图片转成PIL格式,方便pytesseract处理 img_pil = Image.fromarray(img) # 可选:增强图像对比度,提升OCR识别率 enhancer = ImageEnhance.Contrast(img_pil) img_pil = enhancer.enhance(2) # 用pytesseract提取文本 text = pytesseract.image_to_string(img_pil, lang='chi_sim') # 中文识别加lang参数,英文可省略 return text # 测试调用 if __name__ == "__main__": result = get_string(src_path) print("提取到的文本:") print(result)
额外优化建议
- 如果噪点比较多,可以适当调大核的尺寸,比如
np.ones((2,2), np.uint8),增强去噪效果 - 如果图片模糊,可以先加个滤波操作,比如
img = cv2.GaussianBlur(img, (3,3), 0) - 确保已经安装了Tesseract的语言包(比如中文识别需要
chi_sim包)
内容的提问来源于stack exchange,提问作者harmed
相关产品推荐
相关产品推荐

