You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用PDFPlumber提取双栏PDF文本 避免跨列内容混排

双栏PDF文本提取解决方案

问题根本原因

pdfplumber默认的extract_text()方法会将页面内所有文本块按纵向y坐标从上到下、横向x坐标从左到右的规则直接拼接,双栏布局下左右两栏同一高度的文本会被判定为同一行,因此出现跨栏混排的问题。

解决方案1:固定分栏裁剪(适合标准双栏报告)

绝大多数企业年报、ESG报告的双栏为左右等宽的固定布局,直接将页面沿中线裁剪为左右两个区域,按左栏→右栏的顺序提取文本即可解决混排问题。
修改后的代码如下:

import requests
import pdfplumber
from io import BytesIO

def converter(url, split_offset=0):
    text = []
    req = requests.get(url)
    with pdfplumber.open(BytesIO(req.content)) as pdf:
        for page in pdf.pages:
            # 计算双栏分割线x坐标,split_offset可根据排版微调,避免裁到边界文字
            split_x = page.width / 2 + split_offset
            # 分别裁剪左、右栏区域
            left_part = page.crop((0, 0, split_x, page.height))
            right_part = page.crop((split_x, 0, page.width, page.height))
            # 按左→右顺序拼接文本
            page_text = left_part.extract_text() + "\n" + right_part.extract_text()
            text.append(page_text)
    return "\n".join(text)

使用时可以根据实际测试的提取效果调整split_offset数值(单位为像素,一般±10以内即可适配绝大多数布局)。

解决方案2:字符坐标智能排序(适配复杂双栏布局)

如果文档存在跨栏标题、栏宽不一致等复杂布局,可以先获取所有字符的坐标信息,按行分组后分别排序输出:

def converter_complex(url, y_tolerance=3):
    text = []
    req = requests.get(url)
    with pdfplumber.open(BytesIO(req.content)) as pdf:
        for page in pdf.pages:
            # 获取页面所有字符的坐标和内容
            chars = page.chars
            # 按y坐标分组(y_tolerance为行容错值,同一行的字符y坐标差在该范围内会被判定为同一行)
            chars.sort(key=lambda c: (c['top'], c['x0']))
            # 按行拼接
            current_top = None
            current_line = []
            page_lines = []
            for c in chars:
                if current_top is None or abs(c['top'] - current_top) > y_tolerance:
                    if current_line:
                        page_lines.append(''.join(current_line))
                    current_top = c['top']
                    current_line = [c['text']]
                else:
                    current_line.append(c['text'])
            if current_line:
                page_lines.append(''.join(current_line))
            text.append('\n'.join(page_lines))
    return "\n".join(text)

注意事项

  • 上述方法仅适用于可复制的原生PDF,如果是扫描版PDF,需要先通过OCR工具识别文字后再做处理
  • 提取的文本会保留原文的换行和段落结构,可直接用于后续主题建模任务

内容的提问来源于stack exchange,提问作者Ramachandran Ravishankar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 05:24:00