使用Camelot提取多页PDF表格合并至Excel单工作表遇语法错误
问题分析与解决方案
语法错误原因
你遇到的语法错误来自两处:
- 第3页提取代码中,
table_areas列表的最后一个元素后多了一个冗余逗号:' 15, 480, 575, 390',],该冗余逗号会导致Python语法解析失败。 - 手动添加的续行符
\容易因后续缩进或格式问题引发解析异常,建议改用括号自动续行(Python允许括号内参数自动换行,无需手动添加续行符)。
修正后的提取代码
先修正语法错误,确保每页表格能正常提取:
# 提取第1页表格 obstables = camelot.read_pdf( filepath, pages='1', flavor='stream', edge_tol=500, strip_text=' °, kn, m, µbar, mbar, in³, psi,\n', table_areas=[ ' 15, 750, 575, 680', ' 15, 680, 575, 570', ' 15, 570, 575, 460', ' 15, 460, 575, 380', ' 15, 380, 575, 300', ' 15, 300, 575, 240', ' 15, 240, 575, 180', ' 15, 180, 575, 110' ], columns=['','','','','','','',''] ) # 提取第2页表格 obstables1 = camelot.read_pdf( filepath, pages='2', flavor='stream', edge_tol=500, strip_text=' °, kn, m, µbar, mbar, in³, psi,\n', table_areas=[ ' 20, 820, 575, 750', ' 20, 730, 140, 655', ' 20, 635, 270, 560', ' 20, 540, 270, 470' ], columns=['','','',''] ) # 提取第3页表格(已修正多余逗号) obstables2 = camelot.read_pdf( filepath, pages='3', flavor='stream', edge_tol=500, strip_text=' °, kn, m, µbar, mbar, in³, psi,\n', table_areas=[ ' 15, 820, 575, 750', ' 15, 730, 575, 660', ' 15, 640, 575, 570', ' 15, 560, 150, 500', ' 15, 480, 575, 390' # 移除了多余的逗号 ], columns=['','','','',''] )
合并表格并导出到单个Excel工作表
不建议直接操作私有属性_tables,正确做法是合并所有TablesList对象,再将各表格的DataFrame合并为一个大DataFrame后导出:
import pandas as pd # 合并所有TablesList对象 all_tables = obstables + obstables1 + obstables2 # 将所有表格的DataFrame合并为一个大DataFrame(可选添加空行分隔不同表格) combined_df = pd.DataFrame() for table in all_tables: combined_df = pd.concat([combined_df, table.df, pd.DataFrame({'': ['']})], ignore_index=True) # 导出到Excel单个工作表 combined_df.to_excel('merged_tables.xlsx', index=False, header=False)
若不需要空行分隔表格,可简化为:
combined_df = pd.concat([table.df for table in all_tables], ignore_index=True)
内容的提问来源于stack exchange,提问作者Jecook
相关产品推荐
相关产品推荐

