如何用Python提取PDF表格并转为对象?求PyPDF2替代库
可替代的PDF表格提取库推荐
针对PyPDF2无法提取PDF表格的问题,以下是几个实用的替代库,附简单使用示例:
Tabula-py
专门面向PDF表格提取的工具,基于Java版Tabula,能精准识别结构化表格,输出格式支持DataFrame(可轻松转换为自定义类对象),支持指定页面、提取区域等参数。
- 安装:
pip install tabula-py - 示例代码:
import tabula import io def extract_tables_with_tabula(pdf_data): pdf_file = io.BytesIO(pdf_data) # 提取所有页面的表格,返回DataFrame列表 tables = tabula.read_pdf(pdf_file, pages="all", multiple_tables=True) # 自定义表格行类 class TableRow: def __init__(self, row_data): self.col1 = row_data[0] self.col2 = row_data[1] # 根据实际表格字段扩展 # 转换为类对象列表 result = [] for table in tables: for _, row in table.iterrows(): result.append(TableRow(row.tolist())) return result
Camelot
专注于PDF表格提取,对带合并单元格、复杂布局的表格支持更好,提供表格质量评分功能,输出同样为DataFrame格式。
- 安装:
pip install camelot-py[cv](依赖OpenCV,也可选择其他后端) - 示例代码:
import camelot import io def extract_tables_with_camelot(pdf_data): pdf_file = io.BytesIO(pdf_data) # 读取所有页面的表格 tables = camelot.read_pdf(pdf_file, pages="all") # 自定义表格项类 class TableItem: def __init__(self, **kwargs): for k, v in kwargs.items(): setattr(self, k, v) # 转换为类对象列表 result = [] for table in tables: df = table.df for _, row in df.iterrows(): item = TableItem(**row.to_dict()) result.append(item) return result
pdfplumber
基于PyMuPDF构建,提供直观的表格提取API,能识别表格边框、单元格结构,处理半结构化表格表现出色。
- 安装:
pip install pdfplumber - 示例代码:
import pdfplumber import io def extract_tables_with_pdfplumber(pdf_data): pdf_file = io.BytesIO(pdf_data) with pdfplumber.open(pdf_file) as pdf: tables = [] for page in pdf.pages: # 提取当前页面所有表格 page_tables = page.extract_tables() tables.extend(page_tables) # 自定义表格行类 class TableRow: def __init__(self, *cols): self.columns = cols # 转换为类对象列表 return [TableRow(*row) for table in tables for row in table if row]
PyMuPDF(fitz)
轻量高效的PDF处理库,可通过文本块的坐标信息识别表格结构,适合需要自定义表格提取逻辑的场景。
- 安装:
pip install pymupdf - 示例代码(基础行识别):
import fitz import io def extract_tables_with_pymupdf(pdf_data): pdf_file = io.BytesIO(pdf_data) doc = fitz.open(stream=pdf_file, filetype="pdf") all_rows = [] # 自定义表格行类 class TableRow: def __init__(self, row): self.columns = row for page in doc: # 获取页面文本块(包含坐标) blocks = page.get_text("blocks") # 按y坐标分组识别行,x坐标排序识别列 blocks.sort(key=lambda b: (b[1], b[0])) current_row = [] prev_y = None for block in blocks: text = block[4].strip() if not text: continue # 以y坐标差判断是否为同一行(阈值可调整) if prev_y is None or abs(block[1] - prev_y) > 5: if current_row: all_rows.append(TableRow(current_row)) current_row = [text] prev_y = block[1] else: current_row.append(text) if current_row: all_rows.append(TableRow(current_row)) return all_rows
内容的提问来源于stack exchange,提问作者userzxxxxxxx
相关产品推荐
相关产品推荐

