如何用Python从扫描PDF手册中提取表格至Excel?
扫描PDF图片表格提取至Excel的解决方案
问题核心
Camelot、PyMuPDF这类工具仅支持文本型PDF,你的PDF是扫描生成的图片格式,无直接可读取的文本层,因此无法被这类工具识别。必须通过OCR(光学字符识别)技术先将图片中的表格内容转化为结构化文本,再导出至Excel。
方案一:先转可搜索PDF,再复用Camelot处理
先用OCR工具给PDF添加文本层,再用你原有的Camelot代码处理:
- 安装依赖:
pip install ocrmypdf camelot-py[cv] pandas
(需提前安装Tesseract OCR引擎,前往官网下载对应系统版本)
- 生成带文本层的PDF:
import ocrmypdf # 将扫描PDF转换为可搜索PDF,同时矫正倾斜 ocrmypdf.ocr("C:/Users/Vibes/Desktop/Projects/.pdf/Camelot.pdf", "C:/Users/Vibes/Desktop/Projects/.pdf/Camelot_ocr.pdf", deskew=True)
- 用原有代码处理新PDF并导出:
import camelot import pandas as pd file = r"C:/Users/Vibes/Desktop/Projects/.pdf/Camelot_ocr.pdf" tables = camelot.read_pdf(file, pages='all', flavor="stream", encoding="utf-8") master_DF = pd.DataFrame() new_header = None for i in range(tables.n): if i == 0: new_header = tables[i].df.iloc[4] tables[i].df = tables[i].df[5:] tables[i].df.columns = new_header master_DF = pd.concat([master_DF, tables[i].df], axis=0, ignore_index=True) else: tables[i].df = tables[i].df[1:] tables[i].df.columns = new_header master_DF = pd.concat([master_DF, tables[i].df], axis=0, ignore_index=True) # 导出到Excel master_DF.to_excel("extracted_table.xlsx", index=False)
方案二:直接提取PDF图片,用Tesseract识别表格结构
如果方案一精度不足,可直接处理PDF图片并提取表格:
- 安装依赖:
pip install pdf2image pytesseract pandas openpyxl
(需安装Tesseract引擎,未添加环境变量时需在代码中指定路径)
- 代码示例:
from pdf2image import convert_from_path import pytesseract import pandas as pd from pytesseract import Output # 若Tesseract未添加到系统环境变量,取消注释并填写路径 # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' pdf_path = r"C:/Users/Vibes/Desktop/Projects/.pdf/Camelot.pdf" images = convert_from_path(pdf_path) master_DF = pd.DataFrame() for img in images: # 提取结构化OCR数据 ocr_data = pytesseract.image_to_data(img, output_type=Output.DATAFRAME, lang='eng', config='--psm 6 --oem 3') # 过滤无效空文本行 valid_rows = ocr_data[ocr_data['text'].notna() & (ocr_data['text'] != '')] # 按行分组(取整top坐标作为行标识,适配表格行对齐) valid_rows['row'] = valid_rows['top'].round(-1) table_rows = valid_rows.groupby('row')['text'].apply(list).tolist() # 转为DataFrame并合并 if table_rows: temp_df = pd.DataFrame(table_rows[1:], columns=table_rows[0]) master_DF = pd.concat([master_DF, temp_df], ignore_index=True) # 导出到Excel master_DF.to_excel("extracted_table.xlsx", index=False)
注意事项
- Tesseract识别精度依赖PDF清晰度,若存在倾斜、模糊,可先做去噪、矫正预处理。
- 复杂表格可调整Tesseract的
psm参数,比如--psm 4适配列对齐式表格。 - 超复杂表格可尝试商业OCR服务(如Google Cloud Vision、Amazon Textract),精度更高。
内容的提问来源于stack exchange,提问作者Toni
相关产品推荐
相关产品推荐

