如何在For循环中合并多个Pandas DataFrame(PDF Plumber场景)
解决PDF多页数据合并问题
你的代码只输出最后一页处理结果的原因很明确:循环过程中每次都会用新的页数据覆盖temp_pdf变量,最后执行pd.concat([temp_pdf])时,只把最后一页的DataFrame合并进去了,前面所有页的处理结果都被丢弃了。
要合并所有页的处理结果,只需要先初始化一个空列表来存储每一页处理好的DataFrame,循环中把处理完成的temp_pdf添加到列表,最后一次性合并列表里的所有DataFrame即可。
修改后的代码如下:
# 初始化空列表,存储每一页处理后的DataFrame processed_dfs = [] for i in range(len(self.pdf_text)): print(self.pdf_text[i]) temp_pdf = pd.DataFrame(self.pdf_text[i]) temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].apply(lambda x:'(' in str(x)).any()], axis=1,inplace=True) temp_pdf = temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].eq('sky').any()], axis=1) temp_pdf = temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].eq('high').any()], axis=1) temp_pdf = temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].eq('temp').any()], axis=1) temp_pdf = temp_pdf.drop([col for col in temp_pdf.columns if temp_pdf[col].eq('structure)').any()], axis=1) # temp_pdf = temp_pdf.drop(temp_pdf.iloc[:, 4:9], axis=1) temp_pdf.columns = range(temp_pdf.columns.size) # 将当前页处理好的DataFrame加入列表 processed_dfs.append(temp_pdf) # 合并所有页的DataFrame,ignore_index=True重置索引避免重复 combinedpdf = pd.concat(processed_dfs, ignore_index=True) print(combinedpdf)
关键改动说明
- 循环前新增
processed_dfs = []:用来保存每一页处理后的结果,避免被后续循环覆盖 - 循环末尾新增
processed_dfs.append(temp_pdf):把当前页处理好的DataFrame存入列表 - 合并时用
pd.concat(processed_dfs, ignore_index=True):一次性合并列表中所有DataFrame,ignore_index=True会重新生成连续的索引,避免不同页的索引重复导致的问题
内容的提问来源于stack exchange,提问作者user19795989
相关产品推荐
相关产品推荐

