You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pd.concat合并PDF提取表格时报错:无法合并list类型对象

问题描述

我有一个包含100个PDF文件的文件夹,所有PDF的第1页都有需要提取的表格。我想把提取的所有表格合并成一个DataFrame并保存为CSV,但执行pd.concat时报错。


原代码
import os

import camelot
import pandas as pd
import PyPDF2
import tabula

# 设置PDF文件所在的目录路径
dir_path = "my/path/"

# 创建空列表存储表格
tables = []

# 遍历目录中的每个文件
for filename in os.listdir(dir_path):
    # 检查是否为PDF文件
    if filename.endswith(".pdf"):
        # 打开PDF文件
        with open(os.path.join(dir_path, filename), "rb") as pdf_file:
            # 创建PDF阅读器对象
            pdf_reader = PyPDF2.PdfFileReader(pdf_file)

            # 获取PDF的第一页
            page = pdf_reader.getPage(0)

            # 使用tabula-py提取第一页的表格
            table = tabula.read_pdf(pdf_file, pages=1, pandas_options={"header": True})
            print(table)

            # 将表格添加到tables列表中
            tables.append(table)

# 将所有表格合并为单个DataFrame
df = pd.concat(tables)

# 将DataFrame写入CSV文件
df.to_csv("Output.csv", index=False)

报错信息

TypeError: cannot concatenate object of type '<class 'list'>'; only Series and DataFrame objs are valid


问题原因与修复方案

核心原因

tabula.read_pdf()默认返回由DataFrame组成的列表——哪怕每个PDF第一页只有一个表格,返回结果依然是列表格式。你直接把这个列表添加到tables中,导致tables变成了「列表嵌套列表」的结构,pd.concat无法处理这种嵌套类型,因此抛出错误。

另外,代码中引入的PyPDF2完全多余:你用它读取了页面,但tabula会自行处理文件读取逻辑,这部分代码不仅没用,还可能导致文件指针位置异常,影响tabula的读取结果。

修复后的完整代码

import os
import pandas as pd
import tabula

dir_path = "my/path/"
tables = []

for filename in os.listdir(dir_path):
    if filename.endswith(".pdf"):
        file_path = os.path.join(dir_path, filename)
        # 提取列表中的第一个元素(单个DataFrame)
        table = tabula.read_pdf(file_path, pages=1, pandas_options={"header": True})[0]
        tables.append(table)

# 合并时重置索引,避免原索引重复问题
df = pd.concat(tables, ignore_index=True)
df.to_csv("Output.csv", index=False)

额外优化建议

添加异常处理逻辑,避免单个PDF提取失败导致整个脚本中断:

for filename in os.listdir(dir_path):
    if filename.endswith(".pdf"):
        file_path = os.path.join(dir_path, filename)
        try:
            table = tabula.read_pdf(file_path, pages=1, pandas_options={"header": True})[0]
            tables.append(table)
        except Exception as e:
            print(f"处理文件 {filename} 失败: {str(e)}")

内容的提问来源于stack exchange,提问作者akang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 23:55:33