You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整OpenCV代码准确检测扫描PDF表格竖线以实现列分割OCR?

解决OpenCV过度检测表格竖线的问题

需要将扫描PDF文档分割为列以执行OCR和结构化文本提取,当前使用的OpenCV代码出现过度检测,页面被大量红线覆盖,无法准确识别表格竖线。原代码如下:

import cv2
import numpy as np

image = cv2.imread('1968 Plum Book-01.png')

gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
_, binary = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
kernel = np.ones((5, 5), np.uint8)
closed = cv2.morphologyEx(binary, cv2.MORPH_CLOSE, kernel)
lines = cv2.HoughLines(closed, 1, np.pi / 180, 200)

line_image = image.copy()
for line in lines:
    rho, theta = line[0]
    if np.pi / 4 < theta < 3 * np.pi / 4:  # Filter vertical lines (45 to 135 degrees)
        a = np.cos(theta)
        b = np.sin(theta)
        x0 = a * rho
        y0 = b * rho
        x1 = int(x0 + 1000 * (-b))
        y1 = int(y0 + 1000 * (a))
        x2 = int(x0 - 1000 * (-b))
        y2 = int(y0 - 1000 * (a))
        cv2.line(line_image, (x1, y1), (x2, y2), (0, 0, 255), 2)
        
cv2.imshow('Processed Image with Vertical Lines', line_image)
cv2.waitKey(0)
cv2.destroyAllWindows()

以下是针对性的调整方案:

1. 使用竖线专用形态学核

原代码的5x5全1核会同时增强横竖线条,导致横向干扰被误判为竖线。替换为垂直方向的结构元素,只强化竖线特征:

# 竖线专用核:高度远大于宽度,仅增强垂直方向的连续线条
kernel = np.ones((1, 7), np.uint8)  # 可根据表格线粗细调整,比如(1,9)对应更粗的线条
closed = cv2.morphologyEx(binary, cv2.MORPH_CLOSE, kernel)

2. 严格过滤竖线角度并提高霍夫阈值

原代码的角度范围(45-135度)太宽泛,会包含倾斜线;阈值200过低,导致噪声线被检测。调整为:

# 提高阈值减少噪声,缩小theta范围至85-95度(接近纯竖线)
lines = cv2.HoughLines(closed, 1, np.pi / 180, threshold=300)

# 替换角度过滤条件
if (np.pi * 85 / 180) < theta < (np.pi * 95 / 180):

3. 添加竖线去重逻辑

霍夫变换会对同一条竖线检测出多条重合结果,需通过x坐标去重:

line_image = image.copy()
if lines is not None:
    vertical_x = []
    for line in lines:
        rho, theta = line[0]
        if (np.pi * 85 / 180) < theta < (np.pi * 95 / 180):
            # 计算竖线的x坐标(近似值)
            x = int(rho * np.cos(theta))
            vertical_x.append(x)
    
    # 去重:保留间距大于10像素的唯一x坐标(可根据表格列间距调整)
    unique_x = []
    threshold = 10
    vertical_x.sort()
    for x in vertical_x:
        if not unique_x or x - unique_x[-1] > threshold:
            unique_x.append(x)
    
    # 绘制去重后的竖线
    for x in unique_x:
        cv2.line(line_image, (x, 0), (x, image.shape[0]), (0, 0, 255), 2)

4. 优化二值化预处理

添加高斯模糊减少噪声,同时反转二值化结果让表格线变为白色,更适合形态学操作:

gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
# 高斯模糊降噪
blurred = cv2.GaussianBlur(gray, (3, 3), 0)
# 反转二值化,表格线为白色,背景为黑色
_, binary = cv2.threshold(blurred, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)

完整修改后的代码

import cv2
import numpy as np

image = cv2.imread('1968 Plum Book-01.png')

gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
blurred = cv2.GaussianBlur(gray, (3, 3), 0)
_, binary = cv2.threshold(blurred, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)

# 竖线专用形态学核
kernel = np.ones((1, 7), np.uint8)
closed = cv2.morphologyEx(binary, cv2.MORPH_CLOSE, kernel)

# 调整霍夫变换参数
lines = cv2.HoughLines(closed, 1, np.pi / 180, threshold=300)

line_image = image.copy()
if lines is not None:
    vertical_x = []
    for line in lines:
        rho, theta = line[0]
        # 严格过滤竖线角度
        if (np.pi * 85 / 180) < theta < (np.pi * 95 / 180):
            x = int(rho * np.cos(theta))
            vertical_x.append(x)
    
    # 去重处理
    unique_x = []
    threshold = 10
    vertical_x.sort()
    for x in vertical_x:
        if not unique_x or x - unique_x[-1] > threshold:
            unique_x.append(x)
    
    # 绘制唯一竖线
    for x in unique_x:
        cv2.line(line_image, (x, 0), (x, image.shape[0]), (0, 0, 255), 2)

cv2.imshow('Processed Image with Vertical Lines', line_image)
cv2.waitKey(0)
cv2.destroyAllWindows()

内容的提问来源于stack exchange,提问作者Drew Godsell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 12:58:18