如何通过预处理提升Tesseract OCR识别美分浮雕文字Liberty的准确率
提升Tesseract识别林肯硬币"Liberty"的预处理方案
针对你遇到的识别错误问题,以下是几个关键预处理步骤,配合调整后的代码可大幅提升识别成功率:
1. 自适应阈值化增强对比度
硬币表面文字与背景灰度差异小,固定阈值易丢失细节,改用自适应阈值能根据局部区域自动调整,突出文字边缘:
def preprocess_threshold(img): # 自适应高斯阈值,块大小11,常数2,反转黑白让文字为白色 img = cv2.adaptiveThreshold(img, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) return img
2. 形态学操作去除噪声
用开运算(先腐蚀再膨胀)消除硬币表面细小噪点,让文字轮廓更连贯:
def remove_noise(img): kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2,2)) img = cv2.morphologyEx(img, cv2.MORPH_OPEN, kernel, iterations=1) return img
3. 弧形文字展开(核心步骤)
林肯硬币上的"Liberty"是弧形排列的,Tesseract对非水平直线文字识别效果差,需将弧形文字展开为水平:
def unwarp_arc_text(img): h, w = img.shape # 根据硬币布局,假设弧形圆心在图像右侧 center_x = w + int(w * 0.6) center_y = h // 2 # 极坐标转换展开弧形 warped = cv2.warpPolar(img, (w, h), (center_x, center_y), center_x, cv2.WARP_POLAR_LINEAR + cv2.WARP_FILL_OUTLIERS) # 旋转调整为水平方向 warped = cv2.rotate(warped, cv2.ROTATE_90_COUNTERCLOCKWISE) # 裁剪掉无效边缘区域 warped = warped[20:h-20, 20:w-20] return warped
4. 优化Tesseract配置
除--psm 8(单字行识别)外,限制字符集可减少识别错误:
text = pytesseract.image_to_string(processed_img, config='--psm 8 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ')
完整修改后的代码
import cv2 import pytesseract def LoadImage(fn): img = cv2.imread(fn) img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) img = cv2.resize(img, (400,400)) return img def GetLibertyCroppedArea(img): top = 190 bot = 240 lft = 10 rgt = 130 cropped_area = img[top:bot, lft:rgt] cropped_area = cv2.resize(cropped_area, (360,150)) return cropped_area def preprocess_image(img): # 自适应阈值化 img = cv2.adaptiveThreshold(img, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) # 去除噪声 kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2,2)) img = cv2.morphologyEx(img, cv2.MORPH_OPEN, kernel, iterations=1) # 展开弧形文字 img = unwarp_arc_text(img) return img def unwarp_arc_text(img): h, w = img.shape center_x = w + int(w * 0.6) center_y = h // 2 warped = cv2.warpPolar(img, (w, h), (center_x, center_y), center_x, cv2.WARP_POLAR_LINEAR + cv2.WARP_FILL_OUTLIERS) warped = cv2.rotate(warped, cv2.ROTATE_90_COUNTERCLOCKWISE) warped = warped[20:h-20, 20:w-20] return warped # 执行流程 fn = "你的图片路径.jpg" img = LoadImage(fn) liberty = GetLibertyCroppedArea(img) processed_img = preprocess_image(liberty) text = pytesseract.image_to_string(processed_img, config='--psm 8 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ') text = text.strip() print(f"识别结果: {text}")
额外提示
- 若弧形展开效果不佳,可调整
center_x的数值(如w + int(w * 0.5)或w + int(w * 0.7)),找到最适配的圆心位置。 - 可将预处理后的图像放大2-3倍,进一步提升Tesseract的识别精度。
内容的提问来源于stack exchange,提问作者skeeter
相关产品推荐
相关产品推荐

