Tesseract数字识别不准确,如何优化代码提升准确率?
改进Tesseract数字识别准确性的方案
问题背景
有一批仅包含数字的图片,已完成预处理但Tesseract仍无法准确提取数字,典型错误案例:
- 识别结果
0858,预期结果08588(漏识别数字) - 识别结果
95415,预期结果92412(数字混淆) - 识别结果
2043,预期结果20413(漏识别数字) - 识别结果
61416,预期结果61116(数字混淆)
原实现代码:
from PIL import Image, ImageEnhance, ImageFilter import pytesseract CAPTCHA_PATH = 'captcha_images/captcha8.jpeg' RED_REMOVED_PATH = 'processed/red_removed.jpeg' PROCESSED_IMAGE_PATH = 'processed/processed_image.png' def change_pixels_except_black(image_path, output_path, threshold=50): """ Change all pixels to white except for pixels close to black. :param image_path: Path to the input image. :param output_path: Path to save the output image. :param threshold: Threshold to determine if a pixel is black (default is 50). """ # Open the image image = Image.open(image_path) image = image.convert("RGB") # Ensure the image is in RGB mode # Load the image data pixels = image.load() # Get the dimensions of the image width, height = image.size # Iterate over each pixel for y in range(height): for x in range(width): # Get the current pixel's color r, g, b = pixels[x, y] # Check if the pixel is close to black if r < threshold and g < threshold and b < threshold: # Keep the pixel as is (close to black) continue else: # Change the pixel to white pixels[x, y] = (255, 255, 255) # Save the modified image image.save(output_path) def preprocess_image(image_path, output_path): """ Preprocess the image to enhance its quality for OCR. :param image_path: Path to the input image. :param output_path: Path to save the processed image. """ # Open the image image = Image.open(image_path) # Resize the image new_width = image.width * 2 new_height = image.height * 2 image = image.resize((new_width, new_height), Image.LANCZOS) # Convert to grayscale image = image.convert('L') # Increase contrast enhancer = ImageEnhance.Contrast(image) image = enhancer.enhance(2) # Apply a filter to sharpen the image image = image.filter(ImageFilter.SHARPEN) # Save the processed image image.save(output_path) def extract_text(image_path): """ Extract text from an image using pytesseract. :param image_path: Path to the input image. :return: Extracted text. """ text = pytesseract.image_to_string(Image.open( image_path), lang='eng', config='--psm 10 --oem 3 -c tessedit_char_whitelist=0123456789') return text # Example usage change_pixels_except_black(CAPTCHA_PATH, RED_REMOVED_PATH, threshold=70) preprocess_image(RED_REMOVED_PATH, PROCESSED_IMAGE_PATH) text = extract_text(PROCESSED_IMAGE_PATH) print(f"Extracted Text of {CAPTCHA_PATH}:") print(text)
改进方案
1. 替换对比度增强为二值化处理
当前的对比度增强无法彻底分离数字与背景,改用阈值二值化让数字边缘更锐利,减少模糊干扰:
def preprocess_image(image_path, output_path): image = Image.open(image_path) # 放大图片提升分辨率 new_width = image.width * 2 new_height = image.height * 2 image = image.resize((new_width, new_height), Image.LANCZOS) # 转灰度图 image = image.convert('L') # 自适应阈值二值化(适合光照不均的图片) image = image.point(lambda x: 0 if x < 127 else 255, '1') # 若图片光照均匀,可改用固定阈值(根据实际调整阈值) # image = image.point(lambda x: 0 if x < 150 else 255, '1') # 二次锐化强化笔画 image = image.filter(ImageFilter.SHARPEN) image.save(output_path)
2. 调整Tesseract的PSM模式
当前使用--psm 10(单字符识别模式),但图片是多数字序列,应选用适合单行文本的模式:
def extract_text(image_path): # --psm 7:将图片视为单行文本;若确定是固定长度数字,可试--psm 6(单块文本) text = pytesseract.image_to_string(Image.open(image_path), lang='eng', config='--psm 7 --oem 3 -c tessedit_char_whitelist=0123456789') # 清理识别结果中的空格、换行符 text = text.strip() return text
3. 引入OpenCV做形态学降噪
如果图片存在细小噪点或笔画残缺,用OpenCV的形态学操作优化预处理效果:
import cv2 import numpy as np def preprocess_with_opencv(image_path, output_path): # 读取灰度图 img = cv2.imread(image_path, cv2.IMREAD_GRAYSCALE) # 放大图片 img = cv2.resize(img, None, fx=2, fy=2, interpolation=cv2.INTER_LANCZOS4) # 自适应阈值二值化(反转黑白,让数字为白色、背景为黑色) img = cv2.adaptiveThreshold(img, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) # 去除微小噪点 kernel = np.ones((1,1), np.uint8) img = cv2.morphologyEx(img, cv2.MORPH_OPEN, kernel) # 保存处理后的图片 cv2.imwrite(output_path, img)
使用时替换原有的preprocess_image函数即可,OpenCV对验证码类图片的处理效果优于单纯的PIL操作。
4. 针对数字混淆的针对性优化
对于2/5、1/4这类易混淆数字:
- 微调二值化阈值,确保数字笔画完整无断裂
- 若数字长度固定,可在识别后做长度校验,比如预期5位但识别出4位,可检查预处理后的图片是否有漏识别的区域
- 训练Tesseract自定义数字模型:准备标注好的数字样本,生成专属的
.traineddata文件,提升特定字体数字的识别准确率
5. 验证预处理效果
每次调整预处理步骤后,务必查看生成的processed_image.png——如果肉眼都难以分辨数字,Tesseract不可能识别准确,优先保证预处理后的图片数字清晰、无多余噪点、笔画连贯。
内容的提问来源于stack exchange,提问作者Vishwa Mittar
相关产品推荐
相关产品推荐

