如何从TXT文件遍历URL并自动逐个下载批量文件?
解决方案:批量自动逐个下载文件
针对你的需求,推荐两种高效的自动化下载方案,优先选择无需浏览器的requests方案,资源占用更低、速度更快。
方案一:使用requests直接下载(最优选择)
无需启动浏览器,直接通过HTTP请求下载静态文件,适合批量处理大量链接,避免浏览器资源浪费。
import requests from pathlib import Path # 设置文件保存目录,自动创建不存在的文件夹 save_dir = Path("downloaded_files") save_dir.mkdir(exist_ok=True) # 读取链接文件,过滤空行 with open("data_download.txt", "r") as f: urls = [line.strip() for line in f if line.strip()] # 逐个下载文件 for idx, url in enumerate(urls, 1): try: print(f"正在下载第 {idx}/{len(urls)} 个文件: {url}") # 流式下载,避免大文件占用过多内存 response = requests.get(url, stream=True, timeout=30) response.raise_for_status() # 捕获HTTP请求错误 # 从URL中提取文件名 filename = url.split("/")[-1] save_path = save_dir / filename # 分块写入文件 with open(save_path, "wb") as file: for chunk in response.iter_content(chunk_size=8192): file.write(chunk) print(f"下载完成: {save_path}") except Exception as e: print(f"下载失败 {url}: {str(e)}") # 记录失败链接,方便后续重试 with open("failed_urls.txt", "a") as f: f.write(url + "\n") print("所有下载任务完成!")
关键细节
stream=True:流式处理大文件,避免一次性加载到内存timeout=30:设置请求超时时间,防止长时间卡住- 自动创建保存目录,过滤空行避免无效请求
- 记录失败链接,便于后续补下载
方案二:改进Selenium实现自动批量下载
如果必须依赖浏览器(比如需要处理JS渲染的下载逻辑),可通过配置Chrome自动下载,去掉手动点击逻辑,循环处理每个链接。
from selenium import webdriver from selenium.webdriver.chrome.options import Options import time import os # 配置Chrome自动下载参数 chrome_options = Options() download_dir = os.path.abspath("selenium_downloads") os.makedirs(download_dir, exist_ok=True) prefs = { "download.default_directory": download_dir, "download.prompt_for_download": False, "download.directory_upgrade": True, "safebrowsing.enabled": True } chrome_options.add_experimental_option("prefs", prefs) driver = webdriver.Chrome(options=chrome_options) # 读取链接文件 with open("data_download.txt", "r") as f: urls = [line.strip() for line in f if line.strip()] # 检查下载是否完成(通过临时文件后缀判断) def is_download_complete(target_dir): for filename in os.listdir(target_dir): if filename.endswith(".crdownload"): return False return True # 逐个处理链接 for idx, url in enumerate(urls, 1): try: print(f"正在处理第 {idx}/{len(urls)} 个链接: {url}") driver.get(url) # 等待下载完成,最多等待60秒(可根据文件大小调整) wait_time = 0 while not is_download_complete(download_dir) and wait_time < 60: time.sleep(1) wait_time += 1 if wait_time >= 60: print(f"下载超时: {url}") with open("failed_urls.txt", "a") as f: f.write(url + "\n") else: print(f"处理完成: {url}") except Exception as e: print(f"处理失败 {url}: {str(e)}") with open("failed_urls.txt", "a") as f: f.write(url + "\n") driver.quit() print("所有下载任务完成!")
关键细节
- 配置Chrome自动保存到指定目录,无需手动确认下载
- 通过
.crdownload临时文件判断下载是否完成 - 设置超时时间,避免无限等待
- 记录失败链接,便于后续重试
内容的提问来源于stack exchange,提问作者Sagar Rawal
相关产品推荐
相关产品推荐

