You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Camelot批量提取PDF表格时跳过图片格式页面?

解决方案

问题核心是Camelot无法处理图片型PDF页面,且当前代码遇到这类页面时可能因后续逻辑报错导致循环终止。以下是两种可行的解决方式:

方法1:提前检测页面是否为文本型

借助pdfplumber先判断每一页是否包含可提取的文本,仅对文本型页面调用Camelot提取表格:

import camelot
import pdfplumber

def extract_tables_from_pdf(pdf_path):
    tables = []
    with pdfplumber.open(pdf_path) as pdf:
        total_pages = len(pdf.pages)
        for page_num in range(1, total_pages + 1):
            page = pdf.pages[page_num - 1]
            # 检查页面是否有可提取文本
            if page.extract_text():
                try:
                    page_tables = camelot.read_pdf(pdf_path, pages=str(page_num), flavor='lattice')
                    if page_tables:
                        tables.extend(page_tables)
                except Exception as e:
                    print(f"处理第{page_num}页出错: {e}")
            else:
                print(f"跳过图片型页面: 第{page_num}页")
    return tables

# 批量处理PDF循环
pdf_files = ["file1.pdf", "file2.pdf", ...]
for pdf_file in pdf_files:
    print(f"处理文件: {pdf_file}")
    tables = extract_tables_from_pdf(pdf_file)
    # 后续表格处理逻辑(如保存到Excel等)

方法2:捕获警告并跳过无表格页面

如果不想额外引入依赖,可以通过捕获Camelot的警告信息,判断页面类型后继续循环:

import camelot
import warnings

def extract_tables_from_pdf(pdf_path):
    tables = []
    total_pages = camelot.utils.get_page_count(pdf_path)
    for page_num in range(1, total_pages + 1):
        # 捕获所有警告
        with warnings.catch_warnings(record=True) as w:
            warnings.simplefilter("always")
            page_tables = camelot.read_pdf(pdf_path, pages=str(page_num), flavor='lattice')
            # 筛选出图片页相关警告
            has_image_warning = any("image-based" in str(warn.message) for warn in w)
            if has_image_warning:
                print(f"跳过图片型页面: 第{page_num}页")
                continue
            if page_tables:
                tables.extend(page_tables)
            else:
                print(f"第{page_num}页未提取到表格,跳过")
    return tables

# 批量处理PDF循环
pdf_files = ["file1.pdf", "file2.pdf", ...]
for pdf_file in pdf_files:
    print(f"处理文件: {pdf_file}")
    tables = extract_tables_from_pdf(pdf_file)
    # 后续表格处理逻辑

补充说明

  • 两种方案均能跳过图片型页面,保证批量处理循环不会中断。
  • 若需要处理图片型页面中的表格,可额外结合OCR工具(如pytesseract)先将图片转成文本型PDF,再用Camelot处理,但会增加流程复杂度。

内容的提问来源于stack exchange,提问作者redox741

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 02:15:43