You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用glob的For循环运行异常,批量处理PDF转DataFrame失败

代码问题定位与修正方案

原代码的核心问题如下:

  • 循环内生成的DataFrame每次都会被新的计算结果覆盖,没有存储所有PDF的处理结果
  • 函数最终返回空值,调用后无法拿到任何生成的DataFrame
  • 全局的PDF文件列表耦合在函数内部,复用性差
  • 手动调用pdf.close()如果中间出现运行异常,会导致文件句柄泄漏
修正后完整代码
import pdfplumber
import pandas as pd
import glob

# 扫描当前目录下所有PDF文件
pdf_file_list = glob.glob('*.pdf')

def batch_process_pdf(pdf_files):
    # 用字典存储所有处理结果,key为PDF文件名,value为对应DataFrame
    processed_result = {}
    for file_name in pdf_files:
        # 用with上下文管理器自动处理PDF文件关闭,避免资源泄漏
        with pdfplumber.open(file_name) as pdf:
            first_page = pdf.pages[0]
            page_text = first_page.extract_text()
        text_lines = page_text.splitlines()
        df = pd.DataFrame(text_lines, columns=['Location, Billed Amount'])
        df[['Location', 'Billed Amount']] = df['Location, Billed Amount'].str.split('Impressions', n=1, expand=True)
        # 显式指定axis=1适配所有pandas版本,避免语法警告
        df = df.drop('Location, Billed Amount', axis=1)
        # 原代码此处重复给列命名,属于冗余逻辑已删除
        df = df[~df['Billed Amount'].isnull()]
        df = df.replace(' ,', '', regex=True)
        df['Location'] = df['Location'].replace('LA Cluster', 'DTLA', regex=True)
        df['Location'] = df['Location'].replace('\d+', '', regex=True)
        df['Location'] = df['Location'].replace(' ,', '', regex=True)
        df['Location'] = df['Location'].str.strip()
        # 将当前PDF的处理结果存入字典
        processed_result[file_name] = df
    return processed_result

# 执行批量处理,得到所有PDF对应的DataFrame
all_pdf_df = batch_process_pdf(pdf_file_list)
用法说明
  • 处理完成后得到的all_pdf_df是字典结构,键为PDF的原始文件名,值为对应处理好的DataFrame,你可以直接通过文件名取对应数据,比如target_df = all_pdf_df['xxx.pdf']
  • 如果需要遍历所有处理结果,可以用循环实现:
    for pdf_name, df in all_pdf_df.items():
        print(f"文件{pdf_name}的处理结果:")
        print(df)
    

内容的提问来源于stack exchange,提问作者cdlabs45

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 01:24:01