如何在Python中将多页PDF表格数据合并为单个DataFrame
解决多页PDF表格拼接NaN问题的方案
问题核心是第一页的表头行与后续数据页的列结构不匹配,导致pd.concat时列对齐错误产生NaN。以下是两种实用的处理方案:
方案一:分页面读取并统一列名
先单独提取第一页的表头,再给所有数据页指定统一列名后拼接:
import tabula import pandas as pd # 读取第一页表格,提取表头 first_page = tabula.read_pdf('/content/drive/MyDrive/Python_dataset/APEX_Loans_Database_Table (3).pdf', pages=1)[0] column_names = first_page.iloc[0].tolist() # 处理第一页数据:移除表头行,设置正确列名 first_page_data = first_page.iloc[1:].reset_index(drop=True) first_page_data.columns = column_names # 读取剩余所有页面,直接指定列名 remaining_pages = tabula.read_pdf('/content/drive/MyDrive/Python_dataset/APEX_Loans_Database_Table (3).pdf', pages='2-', header=None)[0] remaining_pages.columns = column_names # 拼接所有数据 combined_df = pd.concat([first_page_data, remaining_pages], ignore_index=True) # 可选:清理全空行 combined_df = combined_df.dropna(how='all') print(combined_df)
方案二:批量处理DataFrame列表
如果已经读取了所有页面的DataFrame列表,可批量统一列名并清理无效行:
import tabula import pandas as pd # 读取所有页面的DataFrame列表 dfs = tabula.read_pdf('/content/drive/MyDrive/Python_dataset/APEX_Loans_Database_Table (3).pdf', pages='all') # 从第一页提取表头 header = dfs[0].iloc[0].values.tolist() # 处理第一页:移除表头行,设置列名 processed_dfs = [dfs[0].iloc[1:].reset_index(drop=True)] processed_dfs[0].columns = header # 处理剩余页面:统一设置列名 for df in dfs[1:]: df.columns = header processed_dfs.append(df) # 拼接并清理空行 combined_df = pd.concat(processed_dfs, ignore_index=True) combined_df = combined_df.dropna(how='all') print(combined_df)
额外注意事项
- 如果后续页面列数和表头不一致,需调整tabula参数:比如用
guess=False关闭自动列识别,手动通过columns参数指定列的坐标(可借助tabula的GUI工具获取坐标),确保每一页列数匹配。 - 若第一页没有数据行(仅表头),直接移除第一页的所有行即可,只拼接后续数据页。
内容的提问来源于stack exchange,提问作者Maryum khan
相关产品推荐
相关产品推荐

