You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python处理下载PDF文件时遭遇IndexError问题求助

问题分析与解决:下载等待函数失效导致IndexError

问题根源

你遇到的IndexError本质是download_wait函数没有正确等待文件下载完成,核心原因有两个:

  1. 浏览器创建临时文件的延迟:点击下载按钮后,浏览器需要几毫秒甚至几秒才会生成.crdownload临时文件。原函数在点击后立刻开始检查,此时临时文件还未创建,函数会误判为下载完成直接退出,后续临时文件才生成,导致处理时文件仍处于下载状态。
  2. 等待逻辑的通用性干扰:原函数检查下载文件夹内所有临时文件,若文件夹内存在之前下载残留的.crdownload/.tmp文件,会导致等待超时;反之,若新下载的临时文件还未被遍历到,又会提前结束等待。

从你给出的打印输出也能佐证这点:处理2010年时,文件夹内仍存在NYSE_XOM_2010.pdf.crdownload,说明download_wait没有等到该临时文件消失就返回了。

修复方案

1. 优化下载等待函数,跟踪目标文件

修改等待逻辑,只关注当前年份对应的下载文件,避免无关文件干扰,同时确保等待到目标PDF文件生成且临时文件消失:

import os
import time

def download_wait(target_year):
    download_folder = r"C:\Users\Testuser\Downloads"
    max_wait = 120  # 延长超时时间到2分钟,适配大文件下载
    seconds = 0
    
    while seconds < max_wait:
        time.sleep(1)
        # 检查是否存在当前年份的临时下载文件
        has_temp_file = any(
            (fname.endswith('.crdownload') or fname.endswith('.tmp')) and target_year in fname
            for fname in os.listdir(download_folder)
        )
        # 检查是否存在当前年份的已完成PDF文件
        has_finished_pdf = any(
            fname.lower().endswith('.pdf') and target_year in fname
            for fname in os.listdir(download_folder)
        )
        
        # 两个条件都满足时,说明下载完成
        if has_finished_pdf and not has_temp_file:
            break
        seconds += 1
    return seconds

2. 修复XPATH语法错误

你的XPATH中contains(@href,{year})缺少引号,会导致年份被当作数字解析,而href中的年份通常是字符串,修正后才能正确定位下载按钮:

report = driver.find_elements(By.XPATH, f"//span[@class='btn_archived download'][.//a[contains(@href,'{year}')]]")

3. 增加容错处理,避免索引错误

即使等待函数优化后,仍需对filtered_files为空的情况做处理,防止崩溃:

Years = ["2010", "2011"]
for year in Years:
    try:
        report = driver.find_elements(By.XPATH, f"//span[@class='btn_archived download'][.//a[contains(@href,'{year}')]]")
        if len(report)!= 0:
            report[0].click()
            download_wait(year)
            files = os.listdir(r"C:\Users\Testuser\Downloads")
            # 过滤时只保留当前年份的目标文件
            filtered_files = [
                file for file in files 
                if file.lower().endswith(('.pdf', '.htm')) and year in file
            ]
            print(files, year, filtered_files)
            
            if not filtered_files:
                print(f"未找到{year}对应的有效下载文件,跳过")
                continue
            filename = filtered_files[0]
            # 后续移动文件的逻辑...
    except Exception as e:
        print(f"处理{year}时出错:{str(e)}")
        continue

4. 额外优化建议

  • 避免硬编码下载路径,用变量统一存储,方便后续修改;
  • 每次下载前可清理下载文件夹内的旧文件(若无需保留),避免同名文件干扰;
  • 若下载文件较大,可适当延长max_wait的超时时间。

内容的提问来源于stack exchange,提问作者xxgaryxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 03:16:26