使用Python Selenium批量下载URL文件时提取文件名校验存在性问题
批量URL下载自动校验脚本修改方案
需求背景
要实现一款兼容主流浏览器的自动化批量下载校验脚本,核心需要解决两个问题:
- 自动从任意下载URL中提取对应本地保存的文件名
- 批量触发下载后自动校验文件是否下载成功,最终统计失败率
预期功能逻辑
- 读取存储下载URL列表的文本文件,每行对应一个文件下载地址
- 遍历所有URL,调用浏览器触发下载
- 逐文件校验下载结果,成功打印下载完成提示,失败打印失败提示,最终统计整体失败率
现有代码问题
当前代码仅支持已知固定文件名的单文件下载场景,存在三个明显缺陷:
- 无URL文件名自动提取逻辑,无法适配批量URL场景
- 固定等待5秒的逻辑不可靠:大文件可能5秒未下完,小文件会浪费等待时间
- 无下载结果统计能力
现有代码如下:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import time import os.path import requests op = webdriver.ChromeOptions() p = {'download.default_directory': 'C:\\Users\\VM\\Downloads'} op.add_experimental_option('prefs', p) driver = webdriver.Chrome(executable_path="C:\\chromedriver.exe", options=op) with open("C:\\Users\\VM\\Downloads\\urls.txt",'r') as file: for url in file.readlines(): driver.get(url); time.sleep(5) if os.path.isfile('CHECK URL PATH FILENAME'): print("File download is completed") else: print("File download is not completed")
具体修改实现
核心修改点
- 用Python标准库
urllib.parse解析URL,自动提取路径中的文件名,自动过滤URL携带的查询参数,比手动字符串切割兼容性更强 - 替换固定等待逻辑为轮询校验:循环检查目标文件是否存在、是否存在浏览器临时下载缓存文件(Chrome为
.crdownload后缀、Firefox为.part后缀),设置最长等待超时时间,兼顾下载可靠性和执行效率 - 增加成功/失败计数逻辑,遍历完成后自动计算失败率
- 处理URL读取时的换行符、首尾空白字符问题,避免无效URL触发报错
- 下载路径做全路径拼接,避免相对路径导致的文件校验错误
修改后完整可运行代码
from selenium import webdriver import os import time from urllib.parse import urlparse # 配置项 DOWNLOAD_DIR = "C:\\Users\\VM\\Downloads" CHROMEDRIVER_PATH = "C:\\chromedriver.exe" URL_LIST_PATH = "C:\\Users\\VM\\Downloads\\urls.txt" MAX_WAIT_SECONDS = 30 # 单文件最长等待下载时间 # 初始化浏览器 op = webdriver.ChromeOptions() op.add_experimental_option('prefs', {'download.default_directory': DOWNLOAD_DIR}) driver = webdriver.Chrome(executable_path=CHROMEDRIVER_PATH, options=op) success_count = 0 fail_count = 0 # 读取URL列表 with open(URL_LIST_PATH, 'r', encoding='utf-8') as f: urls = [line.strip() for line in f.readlines() if line.strip()] for url in urls: # 从URL提取文件名 parsed_url = urlparse(url) file_name = os.path.basename(parsed_url.path) # 兼容URL路径为空的异常场景 if not file_name: print(f"URL {url} 无法提取有效文件名,判定下载失败") fail_count +=1 continue file_full_path = os.path.join(DOWNLOAD_DIR, file_name) # 如果本地已存在同名文件,先删除避免误判 if os.path.exists(file_full_path): os.remove(file_full_path) # 触发下载 driver.get(url) # 轮询等待下载完成 is_downloaded = False for _ in range(MAX_WAIT_SECONDS): time.sleep(1) # 检查目标文件存在,且无临时下载缓存文件 if os.path.exists(file_full_path): # 检查是否存在Chrome/Firefox的临时下载文件 temp_file_exists = False for temp_suffix in ['.crdownload', '.part', '.tmp']: if os.path.exists(file_full_path + temp_suffix): temp_file_exists = True break if not temp_file_exists: is_downloaded = True break # 输出结果 if is_downloaded: print(f"文件 {file_name} 下载完成") success_count +=1 else: print(f"文件 {file_name} 下载失败") fail_count +=1 # 输出统计结果 total = success_count + fail_count fail_ratio = (fail_count / total * 100) if total >0 else 0 print(f"\n下载任务完成:总计{total}个文件,成功{success_count}个,失败{fail_count}个,失败率{fail_ratio:.2f}%") driver.quit()
跨浏览器兼容说明
如果需要切换到Firefox、Edge浏览器,只需要替换对应的webdriver初始化、下载目录配置即可,核心的文件名提取、下载校验逻辑完全通用,不需要修改。
内容的提问来源于stack exchange,提问作者tlr
相关产品推荐
相关产品推荐

