You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中高效检测并重命名WebDriver下载的CSV文件?

高效检测CSV下载完成并安全重命名的Python方案

问题核心

你当前的方案依赖.crdownload临时文件检测,但小文件可能直接下载完成不生成该后缀,导致等待逻辑失效;固定等待时长又会浪费时间或适配不了不同大小文件。

优化方案与代码改进

核心思路

结合下载前文件快照、新文件追踪和文件大小稳定检测,同时兼容带临时文件和无临时文件的下载场景,确保文件真正写入磁盘后再执行重命名。

改进后的代码实现

import os
import time
from selenium.common.exceptions import TimeoutException
from selenium.webdriver.support.ui import WebDriverWait

class DownloadHandler:
    def __init__(self, driver, download_dir):
        self.driver = driver
        self.download_dir = download_dir
        self.hf = ...  # 你的点击操作对象
        self.full_college_name = ...  # 你的学院名称变量

    def get_current_files(self):
        # 获取当前下载目录的所有文件完整路径
        return set(os.path.join(self.download_dir, f) for f in os.listdir(self.download_dir))

    def wait_for_new_file(self, initial_files, timeout=300):
        # 等待新文件出现
        start_time = time.time()
        while time.time() - start_time < timeout:
            current_files = self.get_current_files()
            new_files = current_files - initial_files
            if new_files:
                return next(iter(new_files))
            time.sleep(0.5)
        raise TimeoutException("等待新文件超时")

    def wait_for_file_stable(self, file_path, timeout=300, check_interval=1):
        # 等待文件大小连续2次检测一致,确认写入完成
        start_time = time.time()
        prev_size = -1
        stable_count = 0
        while time.time() - start_time < timeout:
            current_size = os.path.getsize(file_path)
            if current_size == prev_size:
                stable_count += 1
                if stable_count >= 2:
                    return
            else:
                stable_count = 0
                prev_size = current_size
            time.sleep(check_interval)
        raise TimeoutException("文件大小稳定超时")

    def download_and_rename(self, prefix):
        # 记录下载前的文件快照,避免误识别旧文件
        initial_files = self.get_current_files()

        # 点击下载按钮
        xpath = '/html/body/app-root/app-layout/section/app-detail/div/div[4]/div[1]/div[2]/button'
        self.hf.click(xpath)
        print('下载按钮已点击')

        try:
            # 等待新文件出现
            new_file = self.wait_for_new_file(initial_files)
            print(f"检测到新文件: {new_file}")

            # 处理带.crdownload后缀的临时文件
            if new_file.endswith('.crdownload'):
                target_file = new_file[:-len('.crdownload')]
                WebDriverWait(self.driver, 300).until(
                    lambda d: os.path.exists(target_file) and not os.path.exists(new_file)
                )
                new_file = target_file

            # 等待文件大小稳定,适配直接完成下载的小文件
            self.wait_for_file_stable(new_file)

            # 执行重命名
            self.rename_file(new_file, prefix)
        except TimeoutException as e:
            print(f"下载流程超时: {str(e)}")
            return

    def rename_file(self, original_file, prefix):
        new_file_name = f"{prefix} {self.full_college_name}.csv"
        new_file_path = os.path.join(self.download_dir, new_file_name)

        # 处理同名文件
        if os.path.exists(new_file_path):
            os.remove(new_file_path)
            print(f"已删除同名文件: {new_file_name}")

        # 执行重命名
        os.rename(original_file, new_file_path)
        print(f"文件已重命名为: {new_file_name}")

关键优化点

  • 精准定位新文件:通过下载前的文件快照,确保只处理本次下载的文件,不会被目录中其他旧文件干扰
  • 兼容两种下载场景:既处理生成.crdownload的大文件,也处理直接完成下载的小文件
  • 确保文件完整性:通过检测文件大小稳定,避免因系统缓存导致的提前重命名
  • 无冗余等待:仅在必要时等待,小文件下载完成后立即执行重命名

额外建议

  • 如果网站提供直接下载链接,可跳过浏览器,用requests库直接下载文件,能更精准控制下载流程(无需依赖浏览器的临时文件机制)
  • 建议将下载目录设置为专属目录,避免其他无关文件干扰检测逻辑

内容的提问来源于stack exchange,提问作者HMS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 22:17:21