如何调整OpenCV代码准确检测扫描PDF表格竖线以实现列分割OCR?
解决OpenCV过度检测表格竖线的问题
需要将扫描PDF文档分割为列以执行OCR和结构化文本提取,当前使用的OpenCV代码出现过度检测,页面被大量红线覆盖,无法准确识别表格竖线。原代码如下:
import cv2 import numpy as np image = cv2.imread('1968 Plum Book-01.png') gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) _, binary = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) kernel = np.ones((5, 5), np.uint8) closed = cv2.morphologyEx(binary, cv2.MORPH_CLOSE, kernel) lines = cv2.HoughLines(closed, 1, np.pi / 180, 200) line_image = image.copy() for line in lines: rho, theta = line[0] if np.pi / 4 < theta < 3 * np.pi / 4: # Filter vertical lines (45 to 135 degrees) a = np.cos(theta) b = np.sin(theta) x0 = a * rho y0 = b * rho x1 = int(x0 + 1000 * (-b)) y1 = int(y0 + 1000 * (a)) x2 = int(x0 - 1000 * (-b)) y2 = int(y0 - 1000 * (a)) cv2.line(line_image, (x1, y1), (x2, y2), (0, 0, 255), 2) cv2.imshow('Processed Image with Vertical Lines', line_image) cv2.waitKey(0) cv2.destroyAllWindows()
以下是针对性的调整方案:
1. 使用竖线专用形态学核
原代码的5x5全1核会同时增强横竖线条,导致横向干扰被误判为竖线。替换为垂直方向的结构元素,只强化竖线特征:
# 竖线专用核:高度远大于宽度,仅增强垂直方向的连续线条 kernel = np.ones((1, 7), np.uint8) # 可根据表格线粗细调整,比如(1,9)对应更粗的线条 closed = cv2.morphologyEx(binary, cv2.MORPH_CLOSE, kernel)
2. 严格过滤竖线角度并提高霍夫阈值
原代码的角度范围(45-135度)太宽泛,会包含倾斜线;阈值200过低,导致噪声线被检测。调整为:
# 提高阈值减少噪声,缩小theta范围至85-95度(接近纯竖线) lines = cv2.HoughLines(closed, 1, np.pi / 180, threshold=300) # 替换角度过滤条件 if (np.pi * 85 / 180) < theta < (np.pi * 95 / 180):
3. 添加竖线去重逻辑
霍夫变换会对同一条竖线检测出多条重合结果,需通过x坐标去重:
line_image = image.copy() if lines is not None: vertical_x = [] for line in lines: rho, theta = line[0] if (np.pi * 85 / 180) < theta < (np.pi * 95 / 180): # 计算竖线的x坐标(近似值) x = int(rho * np.cos(theta)) vertical_x.append(x) # 去重:保留间距大于10像素的唯一x坐标(可根据表格列间距调整) unique_x = [] threshold = 10 vertical_x.sort() for x in vertical_x: if not unique_x or x - unique_x[-1] > threshold: unique_x.append(x) # 绘制去重后的竖线 for x in unique_x: cv2.line(line_image, (x, 0), (x, image.shape[0]), (0, 0, 255), 2)
4. 优化二值化预处理
添加高斯模糊减少噪声,同时反转二值化结果让表格线变为白色,更适合形态学操作:
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) # 高斯模糊降噪 blurred = cv2.GaussianBlur(gray, (3, 3), 0) # 反转二值化,表格线为白色,背景为黑色 _, binary = cv2.threshold(blurred, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
完整修改后的代码
import cv2 import numpy as np image = cv2.imread('1968 Plum Book-01.png') gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) blurred = cv2.GaussianBlur(gray, (3, 3), 0) _, binary = cv2.threshold(blurred, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) # 竖线专用形态学核 kernel = np.ones((1, 7), np.uint8) closed = cv2.morphologyEx(binary, cv2.MORPH_CLOSE, kernel) # 调整霍夫变换参数 lines = cv2.HoughLines(closed, 1, np.pi / 180, threshold=300) line_image = image.copy() if lines is not None: vertical_x = [] for line in lines: rho, theta = line[0] # 严格过滤竖线角度 if (np.pi * 85 / 180) < theta < (np.pi * 95 / 180): x = int(rho * np.cos(theta)) vertical_x.append(x) # 去重处理 unique_x = [] threshold = 10 vertical_x.sort() for x in vertical_x: if not unique_x or x - unique_x[-1] > threshold: unique_x.append(x) # 绘制唯一竖线 for x in unique_x: cv2.line(line_image, (x, 0), (x, image.shape[0]), (0, 0, 255), 2) cv2.imshow('Processed Image with Vertical Lines', line_image) cv2.waitKey(0) cv2.destroyAllWindows()
内容的提问来源于stack exchange,提问作者Drew Godsell
相关产品推荐
相关产品推荐

