如何在Python中高效检测并重命名WebDriver下载的CSV文件?
高效检测CSV下载完成并安全重命名的Python方案
问题核心
你当前的方案依赖.crdownload临时文件检测,但小文件可能直接下载完成不生成该后缀,导致等待逻辑失效;固定等待时长又会浪费时间或适配不了不同大小文件。
优化方案与代码改进
核心思路
结合下载前文件快照、新文件追踪和文件大小稳定检测,同时兼容带临时文件和无临时文件的下载场景,确保文件真正写入磁盘后再执行重命名。
改进后的代码实现
import os import time from selenium.common.exceptions import TimeoutException from selenium.webdriver.support.ui import WebDriverWait class DownloadHandler: def __init__(self, driver, download_dir): self.driver = driver self.download_dir = download_dir self.hf = ... # 你的点击操作对象 self.full_college_name = ... # 你的学院名称变量 def get_current_files(self): # 获取当前下载目录的所有文件完整路径 return set(os.path.join(self.download_dir, f) for f in os.listdir(self.download_dir)) def wait_for_new_file(self, initial_files, timeout=300): # 等待新文件出现 start_time = time.time() while time.time() - start_time < timeout: current_files = self.get_current_files() new_files = current_files - initial_files if new_files: return next(iter(new_files)) time.sleep(0.5) raise TimeoutException("等待新文件超时") def wait_for_file_stable(self, file_path, timeout=300, check_interval=1): # 等待文件大小连续2次检测一致,确认写入完成 start_time = time.time() prev_size = -1 stable_count = 0 while time.time() - start_time < timeout: current_size = os.path.getsize(file_path) if current_size == prev_size: stable_count += 1 if stable_count >= 2: return else: stable_count = 0 prev_size = current_size time.sleep(check_interval) raise TimeoutException("文件大小稳定超时") def download_and_rename(self, prefix): # 记录下载前的文件快照,避免误识别旧文件 initial_files = self.get_current_files() # 点击下载按钮 xpath = '/html/body/app-root/app-layout/section/app-detail/div/div[4]/div[1]/div[2]/button' self.hf.click(xpath) print('下载按钮已点击') try: # 等待新文件出现 new_file = self.wait_for_new_file(initial_files) print(f"检测到新文件: {new_file}") # 处理带.crdownload后缀的临时文件 if new_file.endswith('.crdownload'): target_file = new_file[:-len('.crdownload')] WebDriverWait(self.driver, 300).until( lambda d: os.path.exists(target_file) and not os.path.exists(new_file) ) new_file = target_file # 等待文件大小稳定,适配直接完成下载的小文件 self.wait_for_file_stable(new_file) # 执行重命名 self.rename_file(new_file, prefix) except TimeoutException as e: print(f"下载流程超时: {str(e)}") return def rename_file(self, original_file, prefix): new_file_name = f"{prefix} {self.full_college_name}.csv" new_file_path = os.path.join(self.download_dir, new_file_name) # 处理同名文件 if os.path.exists(new_file_path): os.remove(new_file_path) print(f"已删除同名文件: {new_file_name}") # 执行重命名 os.rename(original_file, new_file_path) print(f"文件已重命名为: {new_file_name}")
关键优化点
- 精准定位新文件:通过下载前的文件快照,确保只处理本次下载的文件,不会被目录中其他旧文件干扰
- 兼容两种下载场景:既处理生成
.crdownload的大文件,也处理直接完成下载的小文件 - 确保文件完整性:通过检测文件大小稳定,避免因系统缓存导致的提前重命名
- 无冗余等待:仅在必要时等待,小文件下载完成后立即执行重命名
额外建议
- 如果网站提供直接下载链接,可跳过浏览器,用
requests库直接下载文件,能更精准控制下载流程(无需依赖浏览器的临时文件机制) - 建议将下载目录设置为专属目录,避免其他无关文件干扰检测逻辑
内容的提问来源于stack exchange,提问作者HMS
相关产品推荐
相关产品推荐

