You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python网页爬虫代码下载速度过慢的问题

解决文件未下载完成就执行后续操作的问题

核心问题是点击下载后,程序没等文件完全落地就执行了shutil.move,导致文件找不到或被占用。下面给你几个实用的解决办法:

方法1:等待文件下载完成(检查存在性+大小稳定)

点击下载后,循环检查目标文件是否存在,并且文件大小不再变化(避免下载过程中文件处于临时状态):

import os
import time
from selenium.webdriver.common.by import By
import shutil

try:
    driver.find_element(By.XPATH,'...').click()
    oldfile = old_path + 'file.xlsx'
    newfile = f'output/FederalExperience_{z1}.xlsx'
    
    # 先等文件出现
    while not os.path.exists(oldfile):
        time.sleep(0.5)
    
    # 再等文件大小稳定(确认下载完成)
    prev_size = -1
    while True:
        curr_size = os.path.getsize(oldfile)
        if curr_size == prev_size:
            break
        prev_size = curr_size
        time.sleep(0.5)
    
    # 现在移动文件就不会出错了
    shutil.move(oldfile, newfile)
except Exception as e:
    print(f"具体错误: {str(e)}")  # 别只打印模糊提示,捕获具体错误更利于排查

方法2:直接指定Chrome下载路径(省去移动步骤)

修改ChromeDriver的配置,让下载的文件直接保存到目标目录,从根源避免移动文件的问题:

import os
import time
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By

# 配置Chrome下载设置
chrome_options = Options()
target_dir = os.path.abspath('output')
prefs = {
    "download.default_directory": target_dir,
    "download.prompt_for_download": False,  # 关闭下载弹窗
    "download.directory_upgrade": True,
    "safebrowsing.enabled": True
}
chrome_options.add_experimental_option("prefs", prefs)

# 初始化带配置的driver
driver = webdriver.Chrome(options=chrome_options)

try:
    driver.find_element(By.XPATH,'...').click()
    temp_file = os.path.join(target_dir, 'file.xlsx')
    final_file = os.path.join(target_dir, f'FederalExperience_{z1}.xlsx')
    
    # 同样等待下载完成
    while not os.path.exists(temp_file):
        time.sleep(0.5)
    
    prev_size = -1
    while True:
        curr_size = os.path.getsize(temp_file)
        if curr_size == prev_size:
            break
        prev_size = curr_size
        time.sleep(0.5)
    
    # 重命名文件即可
    os.rename(temp_file, final_file)
except Exception as e:
    print(f"具体错误: {str(e)}")

额外提醒

  • 别用空except,捕获Exception并打印具体错误,能快速定位是文件不存在、权限问题还是其他原因
  • Mac系统默认下载目录是~/Downloads,确认你的old_path是否指向正确路径
  • 如果网站有反爬限制,可在每次请求后加time.sleep(1),避免因请求太频繁导致下载失败

内容的提问来源于stack exchange,提问作者futur3boy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 09:18:13