You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测扫描PDF文档中文本倾斜角度并实现对齐?

扫描PDF文本倾斜角度检测与对齐解决方案

核心思路

针对印刷体文本和表格类扫描PDF,核心依赖文本行/表格边框的方向检测,同时单独处理180°翻转的特殊场景(此时文本倾斜角为0°但需识别翻转需求)。

基于Python的实现方案(OpenCV + Tesseract OCR)

依赖安装

pip install pdf2image opencv-python pytesseract pillow

代码实现

import cv2
import numpy as np
from pdf2image import convert_from_path
import pytesseract

# 若Tesseract未加入系统PATH,需指定路径
# pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

def detect_rotation_angle(image):
    # 图像预处理:灰度化、降噪、二值化
    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
    blur = cv2.GaussianBlur(gray, (5,5), 0)
    thresh = cv2.threshold(blur, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1]

    # 霍夫变换检测文本行角度
    coords = np.column_stack(np.where(thresh > 0))
    angle = cv2.minAreaRect(coords)[-1]
    # 调整角度至-45°~45°范围
    if angle < -45:
        angle = -(90 + angle)
    else:
        angle = -angle

    # Tesseract OSD检测方向,验证180°翻转情况
    osd = pytesseract.image_to_osd(gray)
    osd_angle = int(osd.split('Orientation in degrees: ')[1].split('\n')[0])
    osd_conf = int(osd.split('Orientation confidence: ')[1].split('\n')[0])

    final_angle = angle
    need_flip_180 = False
    # 优先信任高置信度的OSD 180°检测
    if osd_angle == 180 and osd_conf > 80:
        need_flip_180 = True
    # 倾斜角接近0时,验证是否文本倒置
    elif abs(angle) < 1:
        test_roi = gray[0:100, 0:100]
        test_text = pytesseract.image_to_string(test_roi)
        # 若原区域无法识别有效文本,尝试翻转180°后再检测
        if not test_text.strip():
            flipped_roi = cv2.rotate(test_roi, cv2.ROTATE_180)
            flipped_text = pytesseract.image_to_string(flipped_roi)
            if flipped_text.strip():
                need_flip_180 = True

    return final_angle, need_flip_180

def align_pdf_page(page_image):
    angle, need_flip = detect_rotation_angle(page_image)
    # 旋转校正倾斜
    (h, w) = page_image.shape[:2]
    center = (w // 2, h // 2)
    M = cv2.getRotationMatrix2D(center, angle, 1.0)
    rotated = cv2.warpAffine(page_image, M, (w, h), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE)
    # 处理180°翻转
    if need_flip:
        rotated = cv2.rotate(rotated, cv2.ROTATE_180)
    return rotated

# 批量处理PDF示例
pdf_path = "your_scanned.pdf"
pages = convert_from_path(pdf_path)
for idx, page in enumerate(pages):
    cv_page = cv2.cvtColor(np.array(page), cv2.COLOR_RGB2BGR)
    aligned_page = align_pdf_page(cv_page)
    cv2.imwrite(f"aligned_page_{idx+1}.png", aligned_page)

非编程工具选项

  • Adobe Acrobat Pro:自带「扫描与OCR」功能,自动检测并校正文档旋转与倾斜,适合非技术用户
  • ScanTailor:开源扫描文档处理工具,针对印刷体/表格优化,支持自动倾斜校正与旋转检测
  • gImageReader:结合Tesseract的GUI工具,可手动/自动完成旋转校正

关键注意事项

  • 针对100°-120°这类极端倾斜场景,Tesseract的OSD方向检测比霍夫变换更可靠,核心是识别文本本身的方向
  • 表格类文档可结合边框直线检测辅助验证角度,提升准确率
  • 图像预处理(降噪、二值化)是角度检测准确率的关键前提

内容的提问来源于stack exchange,提问作者Paul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 17:45:24