基于OpenCV的扫描PDF列提取问题:冗余文本移除与多列精准分割
解决方案:扫描PDF列提取与冗余文本移除
1. 移除位置不固定的“KULLANMA”文本
结合OCR文本检测定位目标区域后擦除,无需固定裁剪:
- 先用Tesseract对预处理后的图像做文本检测,获取所有文本的边界框
- 遍历检测结果,筛选出内容为“KULLANMA”的框(需指定土耳其语环境
lang='tur',避免大小写识别偏差) - 用图像背景色(可取框周围平均色或直接用白色)填充该边界区域
- 示例代码片段:
import cv2 import pytesseract from pytesseract import Output img = cv2.imread("scan_page.jpg") # 预处理:转灰度、二值化 gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # 文本检测 d = pytesseract.image_to_data(thresh, output_type=Output.DICT, lang='tur') n_boxes = len(d['text']) for i in range(n_boxes): if d['text'][i].strip().upper() == 'KULLANMA': x, y, w, h = d['left'][i], d['top'][i], d['width'][i], d['height'][i] # 用白色填充目标区域(适配多数扫描件背景) cv2.rectangle(img, (x, y), (x + w, y + h), (255, 255, 255), -1) cv2.imwrite("cleaned_page.jpg", img)
2. 排除页眉区域的错误列识别
两种实用方法规避页眉干扰:
方法1:行投影裁剪页眉
- 计算灰度图的水平投影(每行非零像素数量)
- 找到顶部连续低像素的区域(页眉通常文本稀疏或空白),确定有效内容的起始行
- 裁剪掉页眉区域后再执行列提取
- 示例代码片段:
import numpy as np # 计算水平投影 horizontal_proj = np.sum(thresh, axis=1) # 设定阈值,定位有效内容起始行 header_threshold = 50 # 根据扫描件实际密度调整 header_bottom = 0 for i, val in enumerate(horizontal_proj): if val > header_threshold: header_bottom = i break # 裁剪页眉 cropped_img = img[header_bottom:, :]
方法2:轮廓过滤排除页眉
- 提取列轮廓时,添加过滤条件:只保留垂直位置在页眉区域外、高度达图像一半以上的轮廓
- 示例判断逻辑:
img_h = cropped_img.shape[0] # 假设页眉占顶部15%区域 header_region = img_h * 0.15 contours, _ = cv2.findContours(thresh[header_bottom:, :].copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) for cnt in contours: x, y, w, h = cv2.boundingRect(cnt) # 仅保留非页眉区域、高度足够的轮廓 if h > img_h * 0.5: # 处理列轮廓逻辑 pass
3. 自适应不同列数的PDF列分割
基于垂直投影的空白区域检测实现自动分割:
- 计算预处理后图像的垂直投影(每列非零像素数量)
- 识别连续低像素的列(列间空白分隔线),记录分割位置
- 根据分割位置切割图像,适配3/4/6/8列等任意列数
- 示例代码片段:
img_w = cropped_img.shape[1] gray_cropped = cv2.cvtColor(cropped_img, cv2.COLOR_BGR2GRAY) thresh_cropped = cv2.threshold(gray_cropped, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # 计算垂直投影 vertical_proj = np.sum(thresh_cropped, axis=0) # 识别空白分隔线位置 col_threshold = 20 # 根据扫描件列间距调整 split_positions = [] in_blank = False for i, val in enumerate(vertical_proj): if val < col_threshold and not in_blank: split_positions.append(i) in_blank = True elif val >= col_threshold: in_blank = False # 补充首尾位置,执行分割 split_positions = [0] + split_positions + [img_w] columns = [] for i in range(len(split_positions)-1): start = split_positions[i] end = split_positions[i+1] # 过滤过窄的无效分割 if end - start > img_w * 0.05: col_img = cropped_img[:, start:end] columns.append(col_img)
内容的提问来源于stack exchange,提问作者Zeevac
相关产品推荐
相关产品推荐

