如何用Threadpool加速1000+图片下载?现有脚本耗时过长
使用线程池加速批量图片下载
原脚本是单线程串行下载,处理1000+张图片时效率极低——大部分时间都在等待网络响应。用concurrent.futures.ThreadPoolExecutor实现多线程并行下载,能大幅提升速度,以下是改造后的完整方案:
先修正原脚本的基础错误:
- 模块名应为
requests而非request - 文件写入代码存在缩进错误
改造后的代码:
import requests from concurrent.futures import ThreadPoolExecutor # 图片URL列表 image_urls = [ "http://image_eg_001", "http://image_eg_002", "http://image_eg_003", # 补充更多URL... ] def download_image(img_url): """单个图片的下载逻辑封装""" try: file_name = img_url.split('/')[-1] print(f"Downloading File: {file_name}") # 设置超时时间,避免请求无限挂起 r = requests.get(img_url, stream=True, timeout=10) r.raise_for_status() # 捕获HTTP请求错误(如404、500) with open(file_name, 'wb') as f: # 按块读取写入,节省内存占用 for chunk in r.iter_content(chunk_size=1024): if chunk: f.write(chunk) print(f"Downloaded File: {file_name}") except Exception as e: print(f"Failed to download {img_url}: {str(e)}") if __name__ == "__main__": # 设置线程数,建议10-20(根据目标网站限制调整) with ThreadPoolExecutor(max_workers=15) as executor: # 批量提交所有下载任务 executor.map(download_image, image_urls)
核心注意点:
- 线程数控制:
max_workers不要设置过大,10-20足够。线程过多会触发目标网站反爬机制,或导致本地网络资源耗尽,反而拖慢速度。 - 异常防护:通过
try-except和r.raise_for_status(),保证单张图片下载失败不会中断整个批量任务。 - 内存优化:用
iter_content分块写入,避免一次性加载大图片占用过多内存。
内容的提问来源于stack exchange,提问作者Vincent
相关产品推荐
相关产品推荐

