如何用Python读取指定PDF表格?tabula.io等工具无效,需生成DataFrame
可行的PDF表格提取方案及转pandas DataFrame方法
针对你遇到的常规库提取失效问题,可根据PDF类型(文本型/图像扫描型)选择以下方法:
方法1:OCR+手动规整(适用于扫描图像型PDF)
如果目标PDF是扫描生成的图像文件,常规文本提取工具无法识别内容,需结合OCR工具提取文本后再整理成DataFrame:
依赖工具:
pdf2image(PDF转图像)、pytesseract(Tesseract OCR封装)、opencv-python(图像预处理)步骤:
- 安装依赖:
pip install pdf2image pytesseract pandas opencv-python注意:需单独安装Tesseract本体,Windows从官网下载,Linux用
apt install tesseract-ocr,Mac用brew install tesseract,同时安装斯洛文尼亚语语言包(tesseract-ocr-slv)提升准确率。- 提取并规整代码:
from pdf2image import convert_from_path import pytesseract import pandas as pd import cv2 import numpy as np # PDF转图像 pages = convert_from_path("Tedenski-jedilnik-od-5.pdf") all_text = [] for page in pages: # 转OpenCV格式并二值化预处理 img = cv2.cvtColor(np.array(page), cv2.COLOR_RGB2BGR) gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) _, thresh = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY_INV) # OCR提取文本 text = pytesseract.image_to_string(thresh, lang='slv') all_text.append(text) # 拆分文本为行并转DataFrame rows = [] for page_text in all_text: for line in page_text.split('\n'): stripped_line = line.strip() if stripped_line: cols = [col.strip() for col in stripped_line.split() if col.strip()] rows.append(cols) df = pd.DataFrame(rows) # 可手动添加列名:df.columns = ['周一', '周二', ..., '菜品']
方法2:底层文本提取+坐标分组(适用于复杂结构的文本型PDF)
用pdfminer.six获取文本块的坐标信息,通过坐标判断表格的行和列结构:
- 安装依赖:
pip install pdfminer.six pandas
- 提取代码:
from pdfminer.high_level import extract_pages from pdfminer.layout import LTTextContainer import pandas as pd rows = [] current_row = [] prev_y = None for page_layout in extract_pages("Tedenski-jedilnik-od-5.pdf"): # 按顶部y坐标倒序排序文本块(PDF y轴从下往上) text_blocks = sorted( [elem for elem in page_layout if isinstance(elem, LTTextContainer)], key=lambda x: -x.bbox[1] ) for block in text_blocks: text = block.get_text().strip() if not text: continue curr_y = block.bbox[1] # 按y坐标分组(允许5像素误差) if prev_y is None or abs(curr_y - prev_y) < 5: current_row.append(text) else: rows.append(current_row) current_row = [text] prev_y = curr_y if current_row: rows.append(current_row) df = pd.DataFrame(rows)
方法3:非代码快速方案
如果仅需单次处理,可使用Adobe Acrobat Pro的「导出为Excel」功能,其内置的表格识别引擎对复杂结构兼容性更好,导出后直接用pd.read_excel()读取为DataFrame。
方法4:云端OCR服务(高精度备选)
若本地OCR效果不佳,可使用Google Cloud Vision文档AI或AWS Textract,这类服务能直接识别表格结构并返回结构化数据,需配置API密钥(有免费额度)。以AWS Textract为例:
import boto3 import pandas as pd textract = boto3.client('textract') with open("Tedenski-jedilnik-od-5.pdf", 'rb') as f: response = textract.analyze_document(Document={'Bytes': f.read()}, FeatureTypes=['TABLES']) # 解析表格数据 tables = [] for table_block in [b for b in response['Blocks'] if b['BlockType'] == 'TABLE']: cells = [b for b in response['Blocks'] if 'Relationships' in b and table_block['Id'] in [r['Id'] for r in b['Relationships'] if r['Type'] == 'CHILD']] # 按行排序单元格 cells.sort(key=lambda x: (-x['Geometry']['BoundingBox']['Top'], x['Geometry']['BoundingBox']['Left'])) current_row = [] prev_top = None for cell in cells: # 提取单元格文本 text = ' '.join([response['Blocks'][next(i for i, b in enumerate(response['Blocks']) if b['Id'] == w)]['Text'] for w in cell['Relationships'][0]['Ids']]) curr_top = cell['Geometry']['BoundingBox']['Top'] if prev_top is None or abs(curr_top - prev_top) < 0.01: current_row.append(text) else: tables.append(current_row) current_row = [text] prev_top = curr_top tables.append(current_row) df = pd.DataFrame(tables[0])
内容的提问来源于stack exchange,提问作者Samo
相关产品推荐
相关产品推荐

