You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从扫描PDF手册中提取表格至Excel?

扫描PDF图片表格提取至Excel的解决方案

问题核心

Camelot、PyMuPDF这类工具仅支持文本型PDF,你的PDF是扫描生成的图片格式,无直接可读取的文本层,因此无法被这类工具识别。必须通过OCR(光学字符识别)技术先将图片中的表格内容转化为结构化文本,再导出至Excel。


方案一:先转可搜索PDF,再复用Camelot处理

先用OCR工具给PDF添加文本层,再用你原有的Camelot代码处理:

  1. 安装依赖:
pip install ocrmypdf camelot-py[cv] pandas

(需提前安装Tesseract OCR引擎,前往官网下载对应系统版本)

  1. 生成带文本层的PDF:
import ocrmypdf

# 将扫描PDF转换为可搜索PDF,同时矫正倾斜
ocrmypdf.ocr("C:/Users/Vibes/Desktop/Projects/.pdf/Camelot.pdf", "C:/Users/Vibes/Desktop/Projects/.pdf/Camelot_ocr.pdf", deskew=True)
  1. 用原有代码处理新PDF并导出:
import camelot
import pandas as pd

file = r"C:/Users/Vibes/Desktop/Projects/.pdf/Camelot_ocr.pdf"

tables = camelot.read_pdf(file, pages='all', flavor="stream", encoding="utf-8")

master_DF = pd.DataFrame()
new_header = None

for i in range(tables.n):
    if i == 0:
        new_header = tables[i].df.iloc[4] 
        tables[i].df = tables[i].df[5:]
        tables[i].df.columns = new_header
        master_DF = pd.concat([master_DF, tables[i].df], axis=0, ignore_index=True)
    else:
         tables[i].df = tables[i].df[1:]
         tables[i].df.columns = new_header
         master_DF = pd.concat([master_DF, tables[i].df], axis=0, ignore_index=True)

# 导出到Excel
master_DF.to_excel("extracted_table.xlsx", index=False)

方案二:直接提取PDF图片,用Tesseract识别表格结构

如果方案一精度不足,可直接处理PDF图片并提取表格:

  1. 安装依赖:
pip install pdf2image pytesseract pandas openpyxl

(需安装Tesseract引擎,未添加环境变量时需在代码中指定路径)

  1. 代码示例:
from pdf2image import convert_from_path
import pytesseract
import pandas as pd
from pytesseract import Output

# 若Tesseract未添加到系统环境变量,取消注释并填写路径
# pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

pdf_path = r"C:/Users/Vibes/Desktop/Projects/.pdf/Camelot.pdf"
images = convert_from_path(pdf_path)

master_DF = pd.DataFrame()

for img in images:
    # 提取结构化OCR数据
    ocr_data = pytesseract.image_to_data(img, output_type=Output.DATAFRAME, lang='eng', config='--psm 6 --oem 3')
    
    # 过滤无效空文本行
    valid_rows = ocr_data[ocr_data['text'].notna() & (ocr_data['text'] != '')]
    
    # 按行分组(取整top坐标作为行标识,适配表格行对齐)
    valid_rows['row'] = valid_rows['top'].round(-1)
    table_rows = valid_rows.groupby('row')['text'].apply(list).tolist()
    
    # 转为DataFrame并合并
    if table_rows:
        temp_df = pd.DataFrame(table_rows[1:], columns=table_rows[0])
        master_DF = pd.concat([master_DF, temp_df], ignore_index=True)

# 导出到Excel
master_DF.to_excel("extracted_table.xlsx", index=False)

注意事项

  • Tesseract识别精度依赖PDF清晰度,若存在倾斜、模糊,可先做去噪、矫正预处理。
  • 复杂表格可调整Tesseract的psm参数,比如--psm 4适配列对齐式表格。
  • 超复杂表格可尝试商业OCR服务(如Google Cloud Vision、Amazon Textract),精度更高。

内容的提问来源于stack exchange,提问作者Toni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:05:21