You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中将多页PDF表格数据合并为单个DataFrame

解决多页PDF表格拼接NaN问题的方案

问题核心是第一页的表头行与后续数据页的列结构不匹配,导致pd.concat时列对齐错误产生NaN。以下是两种实用的处理方案:

方案一:分页面读取并统一列名

先单独提取第一页的表头,再给所有数据页指定统一列名后拼接:

import tabula
import pandas as pd

# 读取第一页表格,提取表头
first_page = tabula.read_pdf('/content/drive/MyDrive/Python_dataset/APEX_Loans_Database_Table (3).pdf', pages=1)[0]
column_names = first_page.iloc[0].tolist()

# 处理第一页数据:移除表头行,设置正确列名
first_page_data = first_page.iloc[1:].reset_index(drop=True)
first_page_data.columns = column_names

# 读取剩余所有页面,直接指定列名
remaining_pages = tabula.read_pdf('/content/drive/MyDrive/Python_dataset/APEX_Loans_Database_Table (3).pdf', pages='2-', header=None)[0]
remaining_pages.columns = column_names

# 拼接所有数据
combined_df = pd.concat([first_page_data, remaining_pages], ignore_index=True)
# 可选:清理全空行
combined_df = combined_df.dropna(how='all')

print(combined_df)

方案二:批量处理DataFrame列表

如果已经读取了所有页面的DataFrame列表,可批量统一列名并清理无效行:

import tabula
import pandas as pd

# 读取所有页面的DataFrame列表
dfs = tabula.read_pdf('/content/drive/MyDrive/Python_dataset/APEX_Loans_Database_Table (3).pdf', pages='all')

# 从第一页提取表头
header = dfs[0].iloc[0].values.tolist()

# 处理第一页:移除表头行,设置列名
processed_dfs = [dfs[0].iloc[1:].reset_index(drop=True)]
processed_dfs[0].columns = header

# 处理剩余页面:统一设置列名
for df in dfs[1:]:
    df.columns = header
    processed_dfs.append(df)

# 拼接并清理空行
combined_df = pd.concat(processed_dfs, ignore_index=True)
combined_df = combined_df.dropna(how='all')

print(combined_df)

额外注意事项

  • 如果后续页面列数和表头不一致,需调整tabula参数:比如用guess=False关闭自动列识别,手动通过columns参数指定列的坐标(可借助tabula的GUI工具获取坐标),确保每一页列数匹配。
  • 若第一页没有数据行(仅表头),直接移除第一页的所有行即可,只拼接后续数据页。

内容的提问来源于stack exchange,提问作者Maryum khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 11:24:50