You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取复杂PDF表格遇列标题反向问题求解决方案

解决PDF表格列标题反向提取的问题

一、先处理文本旋转/反向的核心思路

列标题反向通常是因为PDF内的文本本身被旋转了90°或270°,多数提取库会直接读取原始文本方向,所以优先修正文本方向。

1. 用PyMuPDF(fitz)检测并修正旋转文本

PyMuPDF可以识别文本块的旋转标记,提取时直接修正:

import fitz  # 需先安装:pip install pymupdf

def fix_rotated_text_in_pdf(pdf_path, page_idx=0):
    doc = fitz.open(pdf_path)
    page = doc[page_idx]
    blocks = page.get_text("dict")["blocks"]
    output_lines = []

    for block in blocks:
        if "lines" not in block:
            continue
        line_content = []
        for line in block["lines"]:
            for span in line["spans"]:
                # 检查文本旋转标记:16=90°旋转,32=270°旋转
                if span["flags"] & (16 | 32):
                    # 反转字符串恢复正常顺序
                    fixed_text = span["text"][::-1]
                else:
                    fixed_text = span["text"]
                line_content.append(fixed_text)
            output_lines.append(" ".join(line_content))
    doc.close()
    return output_lines

# 调用示例
pdf_file = "your_target.pdf"
fixed_content = fix_rotated_text_in_pdf(pdf_file)
for line in fixed_content:
    print(line)

2. OCR处理扫描版或非文本型PDF

如果是扫描生成的PDF,纯文本提取库无效,需结合OCR工具调整方向后识别:

from pdf2image import convert_from_path
import pytesseract  # 需安装Tesseract OCR引擎和pytesseract包

def ocr_fix_rotated_table(pdf_path, page_idx=0, rotate_angle=270):
    # PDF转单页图片
    pages = convert_from_path(pdf_path, first_page=page_idx+1, last_page=page_idx+1)
    page_img = pages[0]
    # 旋转图片到正确方向
    corrected_img = page_img.rotate(rotate_angle)
    # OCR识别表格内容
    raw_text = pytesseract.image_to_string(corrected_img, config="--psm 6")
    # 拆分并清理行
    return [row.strip() for row in raw_text.split("\n") if row.strip()]

# 调用示例
ocr_result = ocr_fix_rotated_table("your_scanned_pdf.pdf")
for row in ocr_result:
    print(row)

二、适合特殊表格的提取库推荐

Camelot-py

专门针对PDF表格的提取工具,支持处理带旋转、无边框的复杂表格,返回格式为DataFrame,方便后续处理:

import camelot  # 安装:pip install camelot-py[cv]

def extract_table_with_camelot(pdf_path, page_num=1):
    # flavor='stream'适配无边框表格,'lattice'适配有边框表格
    tables = camelot.read_pdf(pdf_path, pages=str(page_num), flavor="stream")
    if not tables:
        return None
    df = tables[0].df
    # 手动修正反向列标题
    df.columns = [col[::-1] for col in df.columns]
    return df

# 调用示例
table_df = extract_table_with_camelot("your_file.pdf")
print(table_df)

Tabula-py

基于Java Tabula的Python封装,适合结构化表格提取,操作简单:

import tabula  # 安装:pip install tabula-py

def extract_table_with_tabula(pdf_path, page_num=1):
    tables = tabula.read_pdf(pdf_path, pages=page_num)
    if not tables:
        return None
    df = tables[0]
    # 修正反向列名
    df.columns = [col[::-1] for col in df.columns]
    return df

# 调用示例
table_df = extract_table_with_tabula("your_file.pdf")
print(table_df)

三、通用反向文本修正技巧

如果提取后仍有部分文本反向,直接用字符串反转即可快速修复:

# 示例:修正反向列标题列表
reversed_headers = ["tcejbuS", "emaneM", "rebmuN"]
fixed_headers = [header[::-1] for header in reversed_headers]
print(fixed_headers)  # 输出:['Subject', 'Name', 'Number']

内容的提问来源于stack exchange,提问作者Rahul Dhir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 08:25:22