使用pd.concat合并PDF提取表格时报错:无法合并list类型对象
问题描述
我有一个包含100个PDF文件的文件夹,所有PDF的第1页都有需要提取的表格。我想把提取的所有表格合并成一个DataFrame并保存为CSV,但执行pd.concat时报错。
原代码
import os import camelot import pandas as pd import PyPDF2 import tabula # 设置PDF文件所在的目录路径 dir_path = "my/path/" # 创建空列表存储表格 tables = [] # 遍历目录中的每个文件 for filename in os.listdir(dir_path): # 检查是否为PDF文件 if filename.endswith(".pdf"): # 打开PDF文件 with open(os.path.join(dir_path, filename), "rb") as pdf_file: # 创建PDF阅读器对象 pdf_reader = PyPDF2.PdfFileReader(pdf_file) # 获取PDF的第一页 page = pdf_reader.getPage(0) # 使用tabula-py提取第一页的表格 table = tabula.read_pdf(pdf_file, pages=1, pandas_options={"header": True}) print(table) # 将表格添加到tables列表中 tables.append(table) # 将所有表格合并为单个DataFrame df = pd.concat(tables) # 将DataFrame写入CSV文件 df.to_csv("Output.csv", index=False)
报错信息
TypeError: cannot concatenate object of type '<class 'list'>'; only Series and DataFrame objs are valid
问题原因与修复方案
核心原因
tabula.read_pdf()默认返回由DataFrame组成的列表——哪怕每个PDF第一页只有一个表格,返回结果依然是列表格式。你直接把这个列表添加到tables中,导致tables变成了「列表嵌套列表」的结构,pd.concat无法处理这种嵌套类型,因此抛出错误。
另外,代码中引入的PyPDF2完全多余:你用它读取了页面,但tabula会自行处理文件读取逻辑,这部分代码不仅没用,还可能导致文件指针位置异常,影响tabula的读取结果。
修复后的完整代码
import os import pandas as pd import tabula dir_path = "my/path/" tables = [] for filename in os.listdir(dir_path): if filename.endswith(".pdf"): file_path = os.path.join(dir_path, filename) # 提取列表中的第一个元素(单个DataFrame) table = tabula.read_pdf(file_path, pages=1, pandas_options={"header": True})[0] tables.append(table) # 合并时重置索引,避免原索引重复问题 df = pd.concat(tables, ignore_index=True) df.to_csv("Output.csv", index=False)
额外优化建议
添加异常处理逻辑,避免单个PDF提取失败导致整个脚本中断:
for filename in os.listdir(dir_path): if filename.endswith(".pdf"): file_path = os.path.join(dir_path, filename) try: table = tabula.read_pdf(file_path, pages=1, pandas_options={"header": True})[0] tables.append(table) except Exception as e: print(f"处理文件 {filename} 失败: {str(e)}")
内容的提问来源于stack exchange,提问作者akang
相关产品推荐
相关产品推荐

