You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从TXT文件遍历URL并自动逐个下载批量文件?

解决方案:批量自动逐个下载文件

针对你的需求,推荐两种高效的自动化下载方案,优先选择无需浏览器的requests方案,资源占用更低、速度更快。

方案一:使用requests直接下载(最优选择)

无需启动浏览器,直接通过HTTP请求下载静态文件,适合批量处理大量链接,避免浏览器资源浪费。

import requests
from pathlib import Path

# 设置文件保存目录,自动创建不存在的文件夹
save_dir = Path("downloaded_files")
save_dir.mkdir(exist_ok=True)

# 读取链接文件,过滤空行
with open("data_download.txt", "r") as f:
    urls = [line.strip() for line in f if line.strip()]

# 逐个下载文件
for idx, url in enumerate(urls, 1):
    try:
        print(f"正在下载第 {idx}/{len(urls)} 个文件: {url}")
        # 流式下载,避免大文件占用过多内存
        response = requests.get(url, stream=True, timeout=30)
        response.raise_for_status()  # 捕获HTTP请求错误

        # 从URL中提取文件名
        filename = url.split("/")[-1]
        save_path = save_dir / filename

        # 分块写入文件
        with open(save_path, "wb") as file:
            for chunk in response.iter_content(chunk_size=8192):
                file.write(chunk)
        print(f"下载完成: {save_path}")

    except Exception as e:
        print(f"下载失败 {url}: {str(e)}")
        # 记录失败链接,方便后续重试
        with open("failed_urls.txt", "a") as f:
            f.write(url + "\n")

print("所有下载任务完成!")

关键细节

  • stream=True:流式处理大文件,避免一次性加载到内存
  • timeout=30:设置请求超时时间,防止长时间卡住
  • 自动创建保存目录,过滤空行避免无效请求
  • 记录失败链接,便于后续补下载

方案二:改进Selenium实现自动批量下载

如果必须依赖浏览器(比如需要处理JS渲染的下载逻辑),可通过配置Chrome自动下载,去掉手动点击逻辑,循环处理每个链接。

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time
import os

# 配置Chrome自动下载参数
chrome_options = Options()
download_dir = os.path.abspath("selenium_downloads")
os.makedirs(download_dir, exist_ok=True)

prefs = {
    "download.default_directory": download_dir,
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
    "safebrowsing.enabled": True
}
chrome_options.add_experimental_option("prefs", prefs)

driver = webdriver.Chrome(options=chrome_options)

# 读取链接文件
with open("data_download.txt", "r") as f:
    urls = [line.strip() for line in f if line.strip()]

# 检查下载是否完成(通过临时文件后缀判断)
def is_download_complete(target_dir):
    for filename in os.listdir(target_dir):
        if filename.endswith(".crdownload"):
            return False
    return True

# 逐个处理链接
for idx, url in enumerate(urls, 1):
    try:
        print(f"正在处理第 {idx}/{len(urls)} 个链接: {url}")
        driver.get(url)
        
        # 等待下载完成,最多等待60秒(可根据文件大小调整)
        wait_time = 0
        while not is_download_complete(download_dir) and wait_time < 60:
            time.sleep(1)
            wait_time += 1
        
        if wait_time >= 60:
            print(f"下载超时: {url}")
            with open("failed_urls.txt", "a") as f:
                f.write(url + "\n")
        else:
            print(f"处理完成: {url}")

    except Exception as e:
        print(f"处理失败 {url}: {str(e)}")
        with open("failed_urls.txt", "a") as f:
            f.write(url + "\n")

driver.quit()
print("所有下载任务完成!")

关键细节

  • 配置Chrome自动保存到指定目录,无需手动确认下载
  • 通过.crdownload临时文件判断下载是否完成
  • 设置超时时间,避免无限等待
  • 记录失败链接,便于后续重试

内容的提问来源于stack exchange,提问作者Sagar Rawal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 08:47:08