You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何缩小验证码字符间距并修正旋转字符以实现OCR识别

解决旋转验证码识别问题:校正字符+紧凑拼接实现Tesseract精准识别

针对你遇到的旋转字符验证码无法被Tesseract识别的问题,这里提供一套完整的解决方案,包含字符旋转校正、间距压缩和OCR识别的全流程代码:

完整实现代码

import cv2
import numpy as np
import pytesseract
from imutils import contours

# 若Tesseract未加入系统环境变量,需手动指定路径
# pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

# 读取验证码图像
image = cv2.imread('input_0.png')
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

# 二值化处理(反转后得到黑字白背景,适配Tesseract识别习惯)
thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU + cv2.THRESH_BINARY_INV)[1]

# 提取字符轮廓并按从左到右排序
cnts = cv2.findContours(thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
cnts = cnts[0] if len(cnts) == 2 else cnts[1]
cnts, _ = contours.sort_contours(cnts, method="left-to-right")

corrected_chars = []
for c in cnts:
    area = cv2.contourArea(c)
    if area > 20:  # 过滤小噪点轮廓
        # 获取字符的最小外接矩形,提取旋转角度
        rect = cv2.minAreaRect(c)
        angle = rect[2]
        
        # 校正旋转角度:minAreaRect返回角度范围是[-90,0),小于-45°时需转换为正向角度
        if angle < -45:
            angle = 90 + angle
        
        # 提取字符ROI并旋转校正
        x, y, w, h = cv2.boundingRect(c)
        roi = thresh[y:y+h, x:x+w]
        (h_roi, w_roi) = roi.shape[:2]
        center = (w_roi // 2, h_roi // 2)
        M = cv2.getRotationMatrix2D(center, angle, 1.0)
        rotated = cv2.warpAffine(roi, M, (w_roi, h_roi), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE)
        
        # 裁剪旋转后的空白区域,让字符更紧凑
        non_zero = np.nonzero(rotated)
        cropped = rotated[np.min(non_zero[0]):np.max(non_zero[0])+1, np.min(non_zero[1]):np.max(non_zero[1])+1]
        corrected_chars.append(cropped)

# 拼接校正后的字符,缩小间距
# 统一所有字符高度,避免排版不齐
max_h = max([char.shape[0] for char in corrected_chars])
resized_chars = []
for char in corrected_chars:
    h, w = char.shape
    scale = max_h / h
    resized = cv2.resize(char, (int(w*scale), max_h), interpolation=cv2.INTER_CUBIC)
    resized_chars.append(resized)

# 紧凑拼接字符,仅留2像素间距
gap = 2
total_w = sum([char.shape[1] for char in resized_chars]) + gap*(len(resized_chars)-1)
final_img = np.ones((max_h, total_w), dtype=np.uint8)*255  # 白色背景

current_x = 0
for char in resized_chars:
    h, w = char.shape
    final_img[0:h, current_x:current_x+w] = char
    current_x += w + gap

# 保存处理后的图像(可选)
cv2.imwrite('final_captcha.png', final_img)

# Tesseract识别配置:限制识别范围为数字+大写字母,指定文本块模式
custom_config = r'--oem 3 --psm 6 -c tessedit_char_whitelist=0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ'
text = pytesseract.image_to_string(final_img, config=custom_config)
print("识别结果:", text.strip())

核心步骤说明

  1. 预处理优化:用THRESH_BINARY_INV反转二值化结果,让字符为黑色、背景为白色,契合Tesseract的识别偏好。
  2. 旋转校正:
    • 通过cv2.minAreaRect获取字符的真实旋转角度,针对垂直方向的字符(角度<-45°)转换校正角度,确保字符完全正立。
    • 旋转后裁剪空白区域,减少无效像素对识别的干扰。
  3. 紧凑拼接:统一字符高度后,用极小间距拼接字符,模拟正常文本的紧凑排版,帮助Tesseract识别连续文本。
  4. OCR参数优化:
    • --psm 6指定输入为单一均匀文本块,避免Tesseract误分割。
    • 用tessedit_char_whitelist限制识别范围,大幅降低错误识别概率。

测试结果

运行代码后,处理后的图像会呈现紧凑排列的正立字符,Tesseract可准确识别出目标文本:784FIK

内容的提问来源于stack exchange,提问作者user16514821

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 22:01:04