You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfplumber提取报纸PDF文章边界框不全的问题求助

解决方案:报纸PDF文章边界框提取

报纸PDF的布局通常是多栏、不规则分块的,依赖表格检测的方法确实不适用,以下是几个更适配的方案:

方案1:基于文本块聚类提取文章边界框

利用pdfplumber提取所有文本单词,通过坐标聚类划分文章区域——同一文章的文本块在同一栏内,x坐标范围相近,且y坐标连续。

import pdfplumber
import numpy as np
from sklearn.cluster import DBSCAN

def get_article_bboxes(page):
    # 提取所有带位置信息的单词
    words = page.extract_words(x_tolerance=2, y_tolerance=2, extra_attrs=["bbox"])
    if not words:
        return []
    
    # 提取所有单词的x坐标(用于分栏聚类)
    x_coords = np.array([(word["bbox"][0] + word["bbox"][2])/2 for word in words]).reshape(-1, 1)
    # 使用DBSCAN聚类分栏(eps控制栏间距阈值)
    dbscan = DBSCAN(eps=50, min_samples=5)
    clusters = dbscan.fit_predict(x_coords)
    
    article_bboxes = []
    # 遍历每个栏(聚类结果)
    for cluster_id in np.unique(clusters):
        if cluster_id == -1:
            continue  # 忽略孤立的噪声点
        # 获取当前栏的所有单词
        cluster_words = [word for idx, word in enumerate(words) if clusters[idx] == cluster_id]
        # 按y坐标排序,确保从上到下
        cluster_words.sort(key=lambda w: w["bbox"][1])
        
        # 合并连续的文本块为文章(按y方向的间隙分组)
        current_group = [cluster_words[0]]
        for word in cluster_words[1:]:
            prev_bbox = current_group[-1]["bbox"]
            # 如果当前单词和上一个单词的y间隙小于阈值,归为同一组
            if word["bbox"][1] - prev_bbox[3] < 15:
                current_group.append(word)
            else:
                # 计算当前组的边界框
                min_x = min(w["bbox"][0] for w in current_group)
                min_y = min(w["bbox"][1] for w in current_group)
                max_x = max(w["bbox"][2] for w in current_group)
                max_y = max(w["bbox"][3] for w in current_group)
                article_bboxes.append((min_x, min_y, max_x, max_y))
                current_group = [word]
        # 处理最后一组
        if current_group:
            min_x = min(w["bbox"][0] for w in current_group)
            min_y = min(w["bbox"][1] for w in current_group)
            max_x = max(w["bbox"][2] for w in current_group)
            max_y = max(w["bbox"][3] for w in current_group)
            article_bboxes.append((min_x, min_y, max_x, max_y))
    
    return article_bboxes

# 测试代码
pdf = pdfplumber.open("2.pdf")
p0 = pdf.pages[0]
article_bboxes = get_article_bboxes(p0)

# 可视化边界框
im = p0.to_image(resolution=150)
for bbox in article_bboxes:
    im.draw_rect(bbox, stroke="red", stroke_width=2)
im.show()

说明:

  • x_tolerance和y_tolerance用于合并相邻的字符为单词,可根据PDF的字体大小调整。
  • DBSCAN的eps参数控制栏与栏之间的最小距离,需根据报纸的栏宽调整。
  • y方向的间隙阈值(15)用于判断是否属于同一文章,可根据行间距调整。

方案2:利用pdfplumber的字符和矩形检测结合

报纸文章通常有明确的栏边界,可通过find_rects()提取页面的分隔线,结合字符分布划分区域:

import pdfplumber

def get_article_bboxes_by_layout(page):
    # 提取页面所有矩形(可能是栏分隔线、边框)
    rects = page.find_rects()
    # 提取所有字符的位置
    chars = page.chars
    if not chars:
        return []
    
    # 提取所有垂直分隔线的x坐标,用于分栏
    vertical_x = sorted(list(set([rect["x0"] for rect in rects if rect["width"] < 5])))
    # 添加页面左右边界
    vertical_x = [0] + vertical_x + [page.width]
    
    article_bboxes = []
    # 遍历每一栏
    for i in range(len(vertical_x)-1):
        col_left = vertical_x[i]
        col_right = vertical_x[i+1]
        # 获取当前栏内的所有字符
        col_chars = [c for c in chars if col_left <= c["x0"] <= col_right]
        if not col_chars:
            continue
        # 按y坐标排序
        col_chars.sort(key=lambda c: c["top"])
        
        # 按y间隙分组为文章
        current_group = [col_chars[0]]
        for char in col_chars[1:]:
            prev_top = current_group[-1]["top"]
            prev_bottom = current_group[-1]["top"] + current_group[-1]["height"]
            # 间隙大于阈值则分为新文章
            if char["top"] - prev_bottom > 20:
                min_x = min(c["x0"] for c in current_group)
                min_y = min(c["top"] for c in current_group)
                max_x = max(c["x1"] for c in current_group)
                max_y = max(c["top"] + c["height"] for c in current_group)
                article_bboxes.append((min_x, min_y, max_x, max_y))
                current_group = [char]
            else:
                current_group.append(char)
        # 处理最后一组
        if current_group:
            min_x = min(c["x0"] for c in current_group)
            min_y = min(c["top"] for c in current_group)
            max_x = max(c["x1"] for c in current_group)
            max_y = max(c["top"] + c["height"] for c in current_group)
            article_bboxes.append((min_x, min_y, max_x, max_y))
    
    return article_bboxes

# 测试代码
pdf = pdfplumber.open("2.pdf")
p0 = pdf.pages[0]
article_bboxes = get_article_bboxes_by_layout(p0)

im = p0.to_image(resolution=150)
for bbox in article_bboxes:
    im.draw_rect(bbox, stroke="blue", stroke_width=2)
im.show()

说明:

  • 假设垂直分隔线的宽度小于5,可根据实际PDF调整该阈值。
  • 若PDF没有明确的分隔线,可跳过矩形检测,直接按固定栏宽划分(比如报纸常见的3-4栏)。

方案3:结合PyMuPDF增强布局分析

PyMuPDF的Page.get_text("dict")能返回更结构化的文本布局信息,可配合pdfplumber使用:

import pdfplumber
import fitz

def get_article_bboxes_with_pymupdf(pdf_path, page_num):
    # 用PyMuPDF提取文本块
    doc = fitz.open(pdf_path)
    page = doc[page_num]
    text_dict = page.get_text("dict")
    blocks = text_dict["blocks"]
    
    article_bboxes = []
    # 过滤掉非文本块,合并同一区域的文本块
    for block in blocks:
        if block["type"] != 0:  # 0为文本块
            continue
        # 文本块的边界框
        bbox = (block["bbox"][0], block["bbox"][1], block["bbox"][2], block["bbox"][3])
        article_bboxes.append(bbox)
    
    # 用pdfplumber可视化
    pdf = pdfplumber.open(pdf_path)
    p0 = pdf.pages[page_num]
    im = p0.to_image(resolution=150)
    for bbox in article_bboxes:
        im.draw_rect(bbox, stroke="green", stroke_width=2)
    im.show()
    return article_bboxes

# 测试代码
get_article_bboxes_with_pymupdf("2.pdf", 0)

说明:

  • PyMuPDF会自动将同一文章/段落的文本归为一个块,适合布局复杂的报纸。
  • 若存在小的文本块(如标题),可额外合并相邻的文本块(通过y坐标间隙判断)。

内容的提问来源于stack exchange,提问作者Rajeev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 14:25:13