You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于OpenCV的扫描PDF列提取问题:冗余文本移除与多列精准分割

解决方案:扫描PDF列提取与冗余文本移除

1. 移除位置不固定的“KULLANMA”文本

结合OCR文本检测定位目标区域后擦除,无需固定裁剪:

  • 先用Tesseract对预处理后的图像做文本检测,获取所有文本的边界框
  • 遍历检测结果,筛选出内容为“KULLANMA”的框(需指定土耳其语环境lang='tur',避免大小写识别偏差)
  • 用图像背景色(可取框周围平均色或直接用白色)填充该边界区域
  • 示例代码片段:
import cv2
import pytesseract
from pytesseract import Output

img = cv2.imread("scan_page.jpg")
# 预处理:转灰度、二值化
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1]

# 文本检测
d = pytesseract.image_to_data(thresh, output_type=Output.DICT, lang='tur')
n_boxes = len(d['text'])
for i in range(n_boxes):
    if d['text'][i].strip().upper() == 'KULLANMA':
        x, y, w, h = d['left'][i], d['top'][i], d['width'][i], d['height'][i]
        # 用白色填充目标区域(适配多数扫描件背景)
        cv2.rectangle(img, (x, y), (x + w, y + h), (255, 255, 255), -1)

cv2.imwrite("cleaned_page.jpg", img)

2. 排除页眉区域的错误列识别

两种实用方法规避页眉干扰:

方法1:行投影裁剪页眉

  • 计算灰度图的水平投影(每行非零像素数量)
  • 找到顶部连续低像素的区域(页眉通常文本稀疏或空白),确定有效内容的起始行
  • 裁剪掉页眉区域后再执行列提取
  • 示例代码片段:
import numpy as np

# 计算水平投影
horizontal_proj = np.sum(thresh, axis=1)
# 设定阈值,定位有效内容起始行
header_threshold = 50  # 根据扫描件实际密度调整
header_bottom = 0
for i, val in enumerate(horizontal_proj):
    if val > header_threshold:
        header_bottom = i
        break
# 裁剪页眉
cropped_img = img[header_bottom:, :]

方法2:轮廓过滤排除页眉

  • 提取列轮廓时,添加过滤条件:只保留垂直位置在页眉区域外、高度达图像一半以上的轮廓
  • 示例判断逻辑:
img_h = cropped_img.shape[0]
# 假设页眉占顶部15%区域
header_region = img_h * 0.15
contours, _ = cv2.findContours(thresh[header_bottom:, :].copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
for cnt in contours:
    x, y, w, h = cv2.boundingRect(cnt)
    # 仅保留非页眉区域、高度足够的轮廓
    if h > img_h * 0.5:
        # 处理列轮廓逻辑
        pass

3. 自适应不同列数的PDF列分割

基于垂直投影的空白区域检测实现自动分割:

  • 计算预处理后图像的垂直投影(每列非零像素数量)
  • 识别连续低像素的列(列间空白分隔线),记录分割位置
  • 根据分割位置切割图像,适配3/4/6/8列等任意列数
  • 示例代码片段:
img_w = cropped_img.shape[1]
gray_cropped = cv2.cvtColor(cropped_img, cv2.COLOR_BGR2GRAY)
thresh_cropped = cv2.threshold(gray_cropped, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1]

# 计算垂直投影
vertical_proj = np.sum(thresh_cropped, axis=0)
# 识别空白分隔线位置
col_threshold = 20  # 根据扫描件列间距调整
split_positions = []
in_blank = False
for i, val in enumerate(vertical_proj):
    if val < col_threshold and not in_blank:
        split_positions.append(i)
        in_blank = True
    elif val >= col_threshold:
        in_blank = False

# 补充首尾位置,执行分割
split_positions = [0] + split_positions + [img_w]
columns = []
for i in range(len(split_positions)-1):
    start = split_positions[i]
    end = split_positions[i+1]
    # 过滤过窄的无效分割
    if end - start > img_w * 0.05:
        col_img = cropped_img[:, start:end]
        columns.append(col_img)

内容的提问来源于stack exchange,提问作者Zeevac

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 05:25:17