You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Threadpool加速1000+图片下载?现有脚本耗时过长

使用线程池加速批量图片下载

原脚本是单线程串行下载,处理1000+张图片时效率极低——大部分时间都在等待网络响应。用concurrent.futures.ThreadPoolExecutor实现多线程并行下载,能大幅提升速度,以下是改造后的完整方案:

先修正原脚本的基础错误:

  • 模块名应为requests而非request
  • 文件写入代码存在缩进错误

改造后的代码:

import requests
from concurrent.futures import ThreadPoolExecutor

# 图片URL列表
image_urls = [
    "http://image_eg_001",
    "http://image_eg_002",
    "http://image_eg_003",
    # 补充更多URL...
]

def download_image(img_url):
    """单个图片的下载逻辑封装"""
    try:
        file_name = img_url.split('/')[-1]
        print(f"Downloading File: {file_name}")
        # 设置超时时间,避免请求无限挂起
        r = requests.get(img_url, stream=True, timeout=10)
        r.raise_for_status()  # 捕获HTTP请求错误(如404、500)
        with open(file_name, 'wb') as f:
            # 按块读取写入,节省内存占用
            for chunk in r.iter_content(chunk_size=1024):
                if chunk:
                    f.write(chunk)
        print(f"Downloaded File: {file_name}")
    except Exception as e:
        print(f"Failed to download {img_url}: {str(e)}")

if __name__ == "__main__":
    # 设置线程数,建议10-20(根据目标网站限制调整)
    with ThreadPoolExecutor(max_workers=15) as executor:
        # 批量提交所有下载任务
        executor.map(download_image, image_urls)

核心注意点:

  • 线程数控制:max_workers不要设置过大,10-20足够。线程过多会触发目标网站反爬机制,或导致本地网络资源耗尽,反而拖慢速度。
  • 异常防护:通过try-except和r.raise_for_status(),保证单张图片下载失败不会中断整个批量任务。
  • 内存优化:用iter_content分块写入,避免一次性加载大图片占用过多内存。

内容的提问来源于stack exchange,提问作者Vincent

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 20:27:26