You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取PDF表格并转为对象?求PyPDF2替代库

可替代的PDF表格提取库推荐

针对PyPDF2无法提取PDF表格的问题,以下是几个实用的替代库,附简单使用示例:

Tabula-py

专门面向PDF表格提取的工具,基于Java版Tabula,能精准识别结构化表格,输出格式支持DataFrame(可轻松转换为自定义类对象),支持指定页面、提取区域等参数。

  • 安装:pip install tabula-py
  • 示例代码:
import tabula
import io

def extract_tables_with_tabula(pdf_data):
    pdf_file = io.BytesIO(pdf_data)
    # 提取所有页面的表格,返回DataFrame列表
    tables = tabula.read_pdf(pdf_file, pages="all", multiple_tables=True)
    
    # 自定义表格行类
    class TableRow:
        def __init__(self, row_data):
            self.col1 = row_data[0]
            self.col2 = row_data[1]
            # 根据实际表格字段扩展
    
    # 转换为类对象列表
    result = []
    for table in tables:
        for _, row in table.iterrows():
            result.append(TableRow(row.tolist()))
    return result

Camelot

专注于PDF表格提取,对带合并单元格、复杂布局的表格支持更好,提供表格质量评分功能,输出同样为DataFrame格式。

  • 安装:pip install camelot-py[cv](依赖OpenCV,也可选择其他后端)
  • 示例代码:
import camelot
import io

def extract_tables_with_camelot(pdf_data):
    pdf_file = io.BytesIO(pdf_data)
    # 读取所有页面的表格
    tables = camelot.read_pdf(pdf_file, pages="all")
    
    # 自定义表格项类
    class TableItem:
        def __init__(self, **kwargs):
            for k, v in kwargs.items():
                setattr(self, k, v)
    
    # 转换为类对象列表
    result = []
    for table in tables:
        df = table.df
        for _, row in df.iterrows():
            item = TableItem(**row.to_dict())
            result.append(item)
    return result

pdfplumber

基于PyMuPDF构建,提供直观的表格提取API,能识别表格边框、单元格结构,处理半结构化表格表现出色。

  • 安装:pip install pdfplumber
  • 示例代码:
import pdfplumber
import io

def extract_tables_with_pdfplumber(pdf_data):
    pdf_file = io.BytesIO(pdf_data)
    with pdfplumber.open(pdf_file) as pdf:
        tables = []
        for page in pdf.pages:
            # 提取当前页面所有表格
            page_tables = page.extract_tables()
            tables.extend(page_tables)
    
    # 自定义表格行类
    class TableRow:
        def __init__(self, *cols):
            self.columns = cols
    
    # 转换为类对象列表
    return [TableRow(*row) for table in tables for row in table if row]

PyMuPDF(fitz)

轻量高效的PDF处理库,可通过文本块的坐标信息识别表格结构,适合需要自定义表格提取逻辑的场景。

  • 安装:pip install pymupdf
  • 示例代码(基础行识别):
import fitz
import io

def extract_tables_with_pymupdf(pdf_data):
    pdf_file = io.BytesIO(pdf_data)
    doc = fitz.open(stream=pdf_file, filetype="pdf")
    all_rows = []
    
    # 自定义表格行类
    class TableRow:
        def __init__(self, row):
            self.columns = row
    
    for page in doc:
        # 获取页面文本块(包含坐标)
        blocks = page.get_text("blocks")
        # 按y坐标分组识别行,x坐标排序识别列
        blocks.sort(key=lambda b: (b[1], b[0]))
        current_row = []
        prev_y = None
        for block in blocks:
            text = block[4].strip()
            if not text:
                continue
            # 以y坐标差判断是否为同一行(阈值可调整)
            if prev_y is None or abs(block[1] - prev_y) > 5:
                if current_row:
                    all_rows.append(TableRow(current_row))
                current_row = [text]
                prev_y = block[1]
            else:
                current_row.append(text)
        if current_row:
            all_rows.append(TableRow(current_row))
    return all_rows

内容的提问来源于stack exchange,提问作者userzxxxxxxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 18:05:30