Python使用glob的For循环运行异常,批量处理PDF转DataFrame失败
代码问题定位与修正方案
原代码的核心问题如下:
- 循环内生成的DataFrame每次都会被新的计算结果覆盖,没有存储所有PDF的处理结果
- 函数最终返回空值,调用后无法拿到任何生成的DataFrame
- 全局的PDF文件列表耦合在函数内部,复用性差
- 手动调用
pdf.close()如果中间出现运行异常,会导致文件句柄泄漏
修正后完整代码
import pdfplumber import pandas as pd import glob # 扫描当前目录下所有PDF文件 pdf_file_list = glob.glob('*.pdf') def batch_process_pdf(pdf_files): # 用字典存储所有处理结果,key为PDF文件名,value为对应DataFrame processed_result = {} for file_name in pdf_files: # 用with上下文管理器自动处理PDF文件关闭,避免资源泄漏 with pdfplumber.open(file_name) as pdf: first_page = pdf.pages[0] page_text = first_page.extract_text() text_lines = page_text.splitlines() df = pd.DataFrame(text_lines, columns=['Location, Billed Amount']) df[['Location', 'Billed Amount']] = df['Location, Billed Amount'].str.split('Impressions', n=1, expand=True) # 显式指定axis=1适配所有pandas版本,避免语法警告 df = df.drop('Location, Billed Amount', axis=1) # 原代码此处重复给列命名,属于冗余逻辑已删除 df = df[~df['Billed Amount'].isnull()] df = df.replace(' ,', '', regex=True) df['Location'] = df['Location'].replace('LA Cluster', 'DTLA', regex=True) df['Location'] = df['Location'].replace('\d+', '', regex=True) df['Location'] = df['Location'].replace(' ,', '', regex=True) df['Location'] = df['Location'].str.strip() # 将当前PDF的处理结果存入字典 processed_result[file_name] = df return processed_result # 执行批量处理,得到所有PDF对应的DataFrame all_pdf_df = batch_process_pdf(pdf_file_list)
用法说明
- 处理完成后得到的
all_pdf_df是字典结构,键为PDF的原始文件名,值为对应处理好的DataFrame,你可以直接通过文件名取对应数据,比如target_df = all_pdf_df['xxx.pdf'] - 如果需要遍历所有处理结果,可以用循环实现:
for pdf_name, df in all_pdf_df.items(): print(f"文件{pdf_name}的处理结果:") print(df)
内容的提问来源于stack exchange,提问作者cdlabs45
相关产品推荐
相关产品推荐

