You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取指定PDF表格?tabula.io等工具无效,需生成DataFrame

可行的PDF表格提取方案及转pandas DataFrame方法

针对你遇到的常规库提取失效问题,可根据PDF类型(文本型/图像扫描型)选择以下方法:

方法1:OCR+手动规整(适用于扫描图像型PDF)

如果目标PDF是扫描生成的图像文件,常规文本提取工具无法识别内容,需结合OCR工具提取文本后再整理成DataFrame:

  • 依赖工具:pdf2image(PDF转图像)、pytesseract(Tesseract OCR封装)、opencv-python(图像预处理)

  • 步骤:

    1. 安装依赖:
    pip install pdf2image pytesseract pandas opencv-python
    

    注意:需单独安装Tesseract本体,Windows从官网下载,Linux用apt install tesseract-ocr,Mac用brew install tesseract,同时安装斯洛文尼亚语语言包(tesseract-ocr-slv)提升准确率。

    1. 提取并规整代码:
    from pdf2image import convert_from_path
    import pytesseract
    import pandas as pd
    import cv2
    import numpy as np
    
    # PDF转图像
    pages = convert_from_path("Tedenski-jedilnik-od-5.pdf")
    all_text = []
    
    for page in pages:
        # 转OpenCV格式并二值化预处理
        img = cv2.cvtColor(np.array(page), cv2.COLOR_RGB2BGR)
        gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
        _, thresh = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY_INV)
        # OCR提取文本
        text = pytesseract.image_to_string(thresh, lang='slv')
        all_text.append(text)
    
    # 拆分文本为行并转DataFrame
    rows = []
    for page_text in all_text:
        for line in page_text.split('\n'):
            stripped_line = line.strip()
            if stripped_line:
                cols = [col.strip() for col in stripped_line.split() if col.strip()]
                rows.append(cols)
    
    df = pd.DataFrame(rows)
    # 可手动添加列名:df.columns = ['周一', '周二', ..., '菜品']
    

方法2:底层文本提取+坐标分组(适用于复杂结构的文本型PDF)

用pdfminer.six获取文本块的坐标信息,通过坐标判断表格的行和列结构:

  • 安装依赖:
pip install pdfminer.six pandas
  • 提取代码:
from pdfminer.high_level import extract_pages
from pdfminer.layout import LTTextContainer
import pandas as pd

rows = []
current_row = []
prev_y = None

for page_layout in extract_pages("Tedenski-jedilnik-od-5.pdf"):
    # 按顶部y坐标倒序排序文本块(PDF y轴从下往上)
    text_blocks = sorted(
        [elem for elem in page_layout if isinstance(elem, LTTextContainer)],
        key=lambda x: -x.bbox[1]
    )
    for block in text_blocks:
        text = block.get_text().strip()
        if not text:
            continue
        curr_y = block.bbox[1]
        # 按y坐标分组(允许5像素误差)
        if prev_y is None or abs(curr_y - prev_y) < 5:
            current_row.append(text)
        else:
            rows.append(current_row)
            current_row = [text]
        prev_y = curr_y
    if current_row:
        rows.append(current_row)

df = pd.DataFrame(rows)

方法3:非代码快速方案

如果仅需单次处理,可使用Adobe Acrobat Pro的「导出为Excel」功能,其内置的表格识别引擎对复杂结构兼容性更好,导出后直接用pd.read_excel()读取为DataFrame。

方法4:云端OCR服务(高精度备选)

若本地OCR效果不佳,可使用Google Cloud Vision文档AI或AWS Textract,这类服务能直接识别表格结构并返回结构化数据,需配置API密钥(有免费额度)。以AWS Textract为例:

import boto3
import pandas as pd

textract = boto3.client('textract')

with open("Tedenski-jedilnik-od-5.pdf", 'rb') as f:
    response = textract.analyze_document(Document={'Bytes': f.read()}, FeatureTypes=['TABLES'])

# 解析表格数据
tables = []
for table_block in [b for b in response['Blocks'] if b['BlockType'] == 'TABLE']:
    cells = [b for b in response['Blocks'] if 'Relationships' in b and table_block['Id'] in [r['Id'] for r in b['Relationships'] if r['Type'] == 'CHILD']]
    # 按行排序单元格
    cells.sort(key=lambda x: (-x['Geometry']['BoundingBox']['Top'], x['Geometry']['BoundingBox']['Left']))
    current_row = []
    prev_top = None
    for cell in cells:
        # 提取单元格文本
        text = ' '.join([response['Blocks'][next(i for i, b in enumerate(response['Blocks']) if b['Id'] == w)]['Text'] for w in cell['Relationships'][0]['Ids']])
        curr_top = cell['Geometry']['BoundingBox']['Top']
        if prev_top is None or abs(curr_top - prev_top) < 0.01:
            current_row.append(text)
        else:
            tables.append(current_row)
            current_row = [text]
        prev_top = curr_top
    tables.append(current_row)

df = pd.DataFrame(tables[0])

内容的提问来源于stack exchange,提问作者Samo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 15:32:46