You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用camelot提取PDF表格遍历TableList时出现TypeError报错如何解决

问题说明

使用camelot提取PDF表格时,需要从返回的TableList中逐个取出单张表格并设置独立变量名,原有代码如下:

tables = camelot.read_pdf("file.pdf", pages = "1")

table = ""
for i in tables:
   globals()['table'+str(i)] = tables[i] 

运行后触发报错:

TypeError: list indices must be integers or slices, not Table

测试场景下PDF第一页共2个表格,实际场景需要处理数百页PDF、共计数十个表格。

报错原因

for i in tables的遍历逻辑中,循环变量i直接拿到的是TableList里存储的Table实例对象,不是整数索引:

  • 用i做tables[i]的索引时,相当于把Table对象当列表下标传入,直接触发类型错误
  • 把Table对象转字符串拼接变量名,生成的变量名无规律,完全不符合预期
修复代码

用enumerate()同时遍历拿到索引和表格对象,索引从1开始计数,匹配日常命名习惯:

import camelot

# 实际使用时pages参数可传"all"读取全文档,也可传"1-200"这类页码范围
tables = camelot.read_pdf("file.pdf", pages="1")

# 批量生成table1、table2...格式的独立变量
for table_index, single_table in enumerate(tables, start=1):
    globals()[f"table{table_index}"] = single_table

运行后测试场景下会自动生成table1、table2两个变量,分别对应第一页的两个表格,直接调用table1.df即可获取对应表格的DataFrame格式数据。

优化建议

如果需要处理的表格总量大,更推荐用字典统一存储所有表格,比批量生成全局变量更易维护,也能避免变量名冲突问题,参考写法:

# 用字典统一存储
table_collect = {}
for table_index, single_table in enumerate(tables, start=1):
    table_collect[f"table{table_index}"] = single_table

# 需要调用时直接通过key取值即可,比如取第一个表格
# print(table_collect["table1"].df)

内容的提问来源于stack exchange,提问作者TomYabo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 10:24:22