You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升PDF手写文本OCR准确率,对标Mac的PEN TO PRINT?

改进手写PDF OCR准确率的思路与实现方案

当前基于OpenCV、pdf2image和Tesseract的手写文本OCR程序存在识别错误率高的问题,以下是针对性的改进思路,目标是达到Mac平台PEN TO PRINT的识别精度,并实现自适应处理能力:

一、自适应图像预处理优化

手写文本的图像质量(背景噪点、笔迹粗细、倾斜度、光照)是影响识别率的核心因素,需替换简单的灰度转换为自适应预处理流程:

  • 自适应二值化:使用cv2.adaptiveThreshold替代全局阈值,自动处理不同区域的光照差异,保留清晰笔迹:
    processed_image = cv2.adaptiveThreshold(gray_img, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
    
  • 动态去噪与笔迹增强:根据图像噪点水平选择滤波方式,再用形态学操作自适应增强笔迹:
    • 通过计算图像方差判断噪点强度,方差小则用高斯模糊,方差大则用中值滤波
    • 基于笔迹轮廓的平均面积动态生成形态学核,对细笔迹用小核膨胀,粗笔迹用大核腐蚀
  • 自动倾斜校正:通过霍夫变换检测文本行的角度,自适应旋转图像校正倾斜:
    coords = np.column_stack(np.where(processed_image > 0))
    angle = cv2.minAreaRect(coords)[-1]
    if angle < -45:
        angle = -(90 + angle)
    else:
        angle = -angle
    (h, w) = processed_image.shape[:2]
    center = (w // 2, h // 2)
    M = cv2.getRotationMatrix2D(center, angle, 1.0)
    processed_image = cv2.warpAffine(processed_image, M, (w, h), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE)
    

二、Tesseract手写识别参数与模型优化

Tesseract默认配置针对印刷体,需调整为手写文本优化模式:

  • 启用手写识别模式:设置PSM(页面分割模式)为适合手写的参数,同时指定手写语言模型:
    custom_config = r'--oem 3 --psm 6 -l eng_hand'  # eng_hand为手写英文模型,中文可用chi_sim_hand
    recognized_text = pytesseract.image_to_string(image, config=custom_config)
    
  • 字符集限制:如果识别的文本有固定字符范围,设置--tessedit_char_whitelist减少干扰:
    custom_config += r' --tessedit_char_whitelist ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789'
    
  • 使用LSTM模型:确保Tesseract版本为4.0+,启用LSTM引擎(--oem 3),相比传统引擎手写识别精度提升明显

三、自适应流程构建

让程序根据输入图像的特征自动选择处理策略:

  • 图像质量评估:计算图像的对比度、亮度、噪点值,自动触发对应的预处理步骤:
    • 对比度低时先执行直方图均衡化(cv2.equalizeHist)或CLAHE增强
    • 检测到纸张纹理噪点时,增加双边滤波(cv2.bilateralFilter)去除纹理保留笔迹
  • 区域自适应处理:用轮廓检测(cv2.findContours)定位文本区域,只对文本区域进行OCR,忽略空白或干扰区域,减少无效识别
  • 多模型融合:对模糊或低质量的手写区域,结合多个预处理参数的结果进行投票,提升识别稳定性

四、修改后的完整代码示例

import cv2
from pdf2image import convert_from_path
import pytesseract
import numpy as np

# Step 1: Convert PDF to images
pdf_path = '/Location of your file/'
images = convert_from_path(pdf_path)  # 补充缺失的图像转换步骤

# Step 2: Adaptive image preprocessing
def preprocess_image(image):
    img_np = np.array(image)
    gray_img = cv2.cvtColor(img_np, cv2.COLOR_RGB2GRAY)
    
    # 自适应二值化
    binary_img = cv2.adaptiveThreshold(gray_img, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
    
    # 自动倾斜校正
    coords = np.column_stack(np.where(binary_img > 0))
    angle = cv2.minAreaRect(coords)[-1]
    if angle < -45:
        angle = -(90 + angle)
    else:
        angle = -angle
    (h, w) = binary_img.shape[:2]
    center = (w // 2, h // 2)
    M = cv2.getRotationMatrix2D(center, angle, 1.0)
    corrected_img = cv2.warpAffine(binary_img, M, (w, h), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE)
    
    # 动态去噪与笔迹增强
    noise_var = np.var(corrected_img)
    if noise_var > 1000:
        denoised_img = cv2.medianBlur(corrected_img, 3)
    else:
        denoised_img = cv2.GaussianBlur(corrected_img, (3, 3), 0)
    
    # 自适应形态学操作
    contours, _ = cv2.findContours(denoised_img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    if contours:
        avg_area = np.mean([cv2.contourArea(c) for c in contours])
        kernel_size = max(1, int(np.sqrt(avg_area) // 5))
        kernel = np.ones((kernel_size, kernel_size), np.uint8)
        enhanced_img = cv2.dilate(denoised_img, kernel, iterations=1)
    else:
        enhanced_img = denoised_img
    
    return enhanced_img

# Step 3: Recognize text with optimized Tesseract config
def recognize_text(image):
    # 使用手写模型+LSTM引擎
    custom_config = r'--oem 3 --psm 6 -l eng_hand --tessedit_char_whitelist ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789'
    recognized_text = pytesseract.image_to_string(image, config=custom_config)
    return recognized_text

# Step 4: Extract text from PDF
def extract_text_from_pdf(images):
    extracted_text = []
    for image in images:
        processed_image = preprocess_image(image)
        text = recognize_text(processed_image)
        extracted_text.append(text)
    return extracted_text

# Step 5: Process the PDF and extract text
extracted_text = extract_text_from_pdf(images)

# Step 6: Print the extracted text
for page_num, text in enumerate(extracted_text):
    print(f"Page {page_num+1}:\n{text}\n")

内容的提问来源于stack exchange,提问作者youssouf toure

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 18:20:18