You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python pdftotext将双栏PDF文本转换为单栏文本

问题描述

我已经用Python的pdftotext工具读取了PDF文本数据,现在需要把这些数据转换成顺序正确的文本,方便从字符串中提取内容——核心需求是将双栏格式的文本转换为单栏格式。

双栏文本示例

With reference to Stone Age, consider the       4.   With reference to Vedic Age, consider the
 following statements:                                following statements:
 1. Microliths are tiny stone artifacts               1. The Aranyakas deal with mysticism,
       belonging to Middle Stone Age.                       rites, rituals and sacrifices.
 2. The use of bow and arrow began during             2. Child marriage and practice of sati was
       the Old Stone Age                                    prevelant during the Rig Vedic Period.
 3. Lakhudiyar caves of Uttrakhand bear               3. Nishka,Satamana and Krishnala were
       the famous pre-historic cave paintings               types of coins used as medium of
       of wavy lines and hand-linked dancing                exchange.
       figures                                        Which of the statements given above are
                                                      correct?
 Which of the statements given above are
                                                      (a) 1 and 2 only
 correct?
                                                      (b) 2 and 3 only
 (a) 1 and 2 only
                                                      (c) 1 and 3 only
 (b) 2 and 3 only
                                                      (d) 1,2 and 3
 (c) 1 and 3 only
 (d) 1, 2 and 3

当前读取PDF的代码

def extract_text_from_pdf(pdf_path):
    text = ""
    # Load your PDF
    with open(pdf_path, "rb") as f:
        pdf = pdftotext.PDF(f)
    return pdf

解决思路

双栏转单栏的核心是按阅读顺序重组内容:先完整提取左栏所有内容,再提取右栏内容。以下是两种实用实现方案:

方法1:基于文本位置的精准拆分(推荐)

pdftotext仅提取纯文本,不保留位置信息。要精准识别栏位,建议使用能获取文本坐标的pdfplumber库,通过判断文本所在的x坐标区分左右栏:

import pdfplumber

def extract_two_column_text(pdf_path):
    single_column_text = []
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            page_width = page.width
            # 取页面中线作为左右栏拆分阈值,可根据实际排版调整
            split_threshold = page_width / 2
            
            lines = page.extract_text_lines()
            left_column = []
            right_column = []
            
            for line in lines:
                # 根据文本起始x坐标判断所属栏位
                if line["x0"] < split_threshold:
                    left_column.append(line["text"])
                else:
                    right_column.append(line["text"])
            
            # 按左栏→右栏的顺序合并
            single_column_text.extend(left_column)
            single_column_text.extend(right_column)
    
    return "\n".join(single_column_text)

该方案适合排版规范的双栏PDF,精准度高;若栏位拆分不是严格居中,可手动调整split_threshold的值。

方法2:基于内容规则的拆分(适配纯文本场景)

如果只能使用pdftotext提取的纯文本,可通过内容特征(如题号、多空格分隔符)拆分左右栏:

import pdftotext
import re

def convert_two_column_to_single(pdf_path):
    with open(pdf_path, "rb") as f:
        pdf = pdftotext.PDF(f)
    
    single_column = []
    left_buffer = []
    right_buffer = []
    
    for page in pdf:
        lines = page.split("\n")
        for line in lines:
            # 匹配右栏起始的题号特征(如"4.   ")
            match = re.search(r"(\d+\.\s{2,})", line)
            if match:
                split_idx = match.start()
                left_part = line[:split_idx].strip()
                right_part = line[split_idx:].strip()
                if left_part:
                    left_buffer.append(left_part)
                if right_part:
                    right_buffer.append(right_part)
            else:
                # 用多空格拆分左右栏内容
                parts = line.split("       ")
                if len(parts) == 2:
                    left_buffer.append(parts[0].strip())
                    right_buffer.append(parts[1].strip())
                else:
                    # 无法拆分的行默认归入左栏,可根据实际调整
                    left_buffer.append(line.strip())
        
        # 合并当前页面的左右栏内容
        single_column.extend(left_buffer)
        single_column.extend(right_buffer)
        left_buffer.clear()
        right_buffer.clear()
    
    return "\n".join(single_column)

该方案依赖PDF的排版规则,若文本格式不统一,需调整正则表达式或拆分逻辑。

注意事项

  • 若PDF是扫描件,pdftotext无法提取文本,需先用pytesseract做OCR处理,再结合上述方案拆分。
  • 不同PDF的双栏排版可能存在差异,需根据实际样本调整拆分规则或坐标阈值。

内容的提问来源于stack exchange,提问作者Granth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 03:35:21