You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Camelot提取多页PDF表格合并至Excel单工作表遇语法错误

问题分析与解决方案

语法错误原因

你遇到的语法错误来自两处:

  1. 第3页提取代码中,table_areas列表的最后一个元素后多了一个冗余逗号:' 15, 480, 575, 390',],该冗余逗号会导致Python语法解析失败。
  2. 手动添加的续行符\容易因后续缩进或格式问题引发解析异常,建议改用括号自动续行(Python允许括号内参数自动换行,无需手动添加续行符)。

修正后的提取代码

先修正语法错误,确保每页表格能正常提取:

# 提取第1页表格
obstables = camelot.read_pdf(
    filepath,
    pages='1',
    flavor='stream',
    edge_tol=500,
    strip_text=' °, kn, m, µbar, mbar, in³, psi,\n',
    table_areas=[
        ' 15, 750, 575, 680',
        ' 15, 680, 575, 570',
        ' 15, 570, 575, 460',
        ' 15, 460, 575, 380',
        ' 15, 380, 575, 300',
        ' 15, 300, 575, 240',
        ' 15, 240, 575, 180',
        ' 15, 180, 575, 110'
    ],
    columns=['','','','','','','','']
)

# 提取第2页表格
obstables1 = camelot.read_pdf(
    filepath,
    pages='2',
    flavor='stream',
    edge_tol=500,
    strip_text=' °, kn, m, µbar, mbar, in³, psi,\n',
    table_areas=[
        ' 20, 820, 575, 750',
        ' 20, 730, 140, 655',
        ' 20, 635, 270, 560',
        ' 20, 540, 270, 470'
    ],
    columns=['','','','']
)

# 提取第3页表格(已修正多余逗号)
obstables2 = camelot.read_pdf(
    filepath,
    pages='3',
    flavor='stream',
    edge_tol=500,
    strip_text=' °, kn, m, µbar, mbar, in³, psi,\n',
    table_areas=[
        ' 15, 820, 575, 750',
        ' 15, 730, 575, 660',
        ' 15, 640, 575, 570',
        ' 15, 560, 150, 500',
        ' 15, 480, 575, 390'  # 移除了多余的逗号
    ],
    columns=['','','','','']
)

合并表格并导出到单个Excel工作表

不建议直接操作私有属性_tables,正确做法是合并所有TablesList对象,再将各表格的DataFrame合并为一个大DataFrame后导出:

import pandas as pd

# 合并所有TablesList对象
all_tables = obstables + obstables1 + obstables2

# 将所有表格的DataFrame合并为一个大DataFrame(可选添加空行分隔不同表格)
combined_df = pd.DataFrame()
for table in all_tables:
    combined_df = pd.concat([combined_df, table.df, pd.DataFrame({'': ['']})], ignore_index=True)

# 导出到Excel单个工作表
combined_df.to_excel('merged_tables.xlsx', index=False, header=False)

若不需要空行分隔表格,可简化为:

combined_df = pd.concat([table.df for table in all_tables], ignore_index=True)

内容的提问来源于stack exchange,提问作者Jecook

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 11:36:16