You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在For循环中合并多个Pandas DataFrame(PDF Plumber场景)

解决PDF多页数据合并问题

你的代码只输出最后一页处理结果的原因很明确:循环过程中每次都会用新的页数据覆盖temp_pdf变量,最后执行pd.concat([temp_pdf])时,只把最后一页的DataFrame合并进去了,前面所有页的处理结果都被丢弃了。

要合并所有页的处理结果,只需要先初始化一个空列表来存储每一页处理好的DataFrame,循环中把处理完成的temp_pdf添加到列表,最后一次性合并列表里的所有DataFrame即可。

修改后的代码如下:

# 初始化空列表,存储每一页处理后的DataFrame
processed_dfs = []

for i in range(len(self.pdf_text)):
    print(self.pdf_text[i])

    temp_pdf = pd.DataFrame(self.pdf_text[i])
    temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].apply(lambda x:'(' in str(x)).any()], axis=1,inplace=True)
    temp_pdf = temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].eq('sky').any()], axis=1)
    temp_pdf = temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].eq('high').any()], axis=1)
    temp_pdf = temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].eq('temp').any()], axis=1)
    temp_pdf = temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].eq('structure)').any()], axis=1)
    # temp_pdf = temp_pdf.drop(temp_pdf.iloc[:, 4:9], axis=1)
    temp_pdf.columns = range(temp_pdf.columns.size)
    
    # 将当前页处理好的DataFrame加入列表
    processed_dfs.append(temp_pdf)

# 合并所有页的DataFrame,ignore_index=True重置索引避免重复
combinedpdf = pd.concat(processed_dfs, ignore_index=True)
print(combinedpdf)

关键改动说明

  • 循环前新增processed_dfs = []:用来保存每一页处理后的结果,避免被后续循环覆盖
  • 循环末尾新增processed_dfs.append(temp_pdf):把当前页处理好的DataFrame存入列表
  • 合并时用pd.concat(processed_dfs, ignore_index=True):一次性合并列表中所有DataFrame,ignore_index=True会重新生成连续的索引,避免不同页的索引重复导致的问题

内容的提问来源于stack exchange,提问作者user19795989

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 12:43:29