如何基于字典列表实现多文件多线程下载?含GIL相关疑问
多文件并发下载解决方案
先澄清GIL的误区
Python的GIL确实会限制CPU密集型任务的多线程性能,但下载属于IO密集型任务——在等待网络响应时,线程会主动释放GIL,所以多线程完全适合这类场景,能有效提升并发效率,不用被GIL的说法束缚。
方案1:用ThreadPoolExecutor实现多线程下载(推荐)
concurrent.futures.ThreadPoolExecutor是Python标准库中更简洁的并发工具,不需要手动管理线程和信号量,直接指定最大并发数即可。
步骤说明
- 先把原始字典列表解析成扁平的任务元组(目标路径+单个URL),方便批量处理
- 定义下载函数,处理路径创建和文件下载逻辑
- 用线程池控制最多10个并发任务
示例代码
import wget from pathlib import Path from concurrent.futures import ThreadPoolExecutor # 你的原始任务列表 testList = [ {'/test/doc/': ['http://firstfile.tar', 'http://secondfile.tar']}, {'/test/doc/file1': ['http://ulr/somefile.zip', 'http://another file.zip']}, {'/test1': 'http://newfile.tar'} ] def parse_tasks(task_list): """把嵌套的任务列表转换成扁平的(目标路径, URL)元组列表""" tasks = [] for item in task_list: for dest_path, urls in item.items(): # 兼容单个URL的字符串格式 urls = [urls] if isinstance(urls, str) else urls for url in urls: tasks.append((dest_path.strip(), url.strip())) return tasks def download_task(dest_path, url): """单个文件的下载逻辑""" try: dest = Path(dest_path) # 如果目标是文件夹,先创建目录,用URL中的文件名保存 if dest.suffix == '': dest.mkdir(parents=True, exist_ok=True) wget.download(url, out=str(dest)) # 如果目标是具体文件路径,直接指定输出 else: dest.parent.mkdir(parents=True, exist_ok=True) wget.download(url, out=str(dest)) print(f"\n下载完成: {url}") except Exception as e: print(f"\n下载失败 {url}: {str(e)}") if __name__ == "__main__": max_concurrent = 10 tasks = parse_tasks(testList) # 启动线程池执行任务 with ThreadPoolExecutor(max_workers=max_concurrent) as executor: executor.map(lambda x: download_task(*x), tasks) print("所有下载任务完成")
方案2:手动用threading+信号量控制并发
如果需要更精细的线程管理,可以手动创建线程,用Semaphore限制最大并发数。
示例代码
import threading import wget from pathlib import Path testList = [ {'/test/doc/': ['http://firstfile.tar', 'http://secondfile.tar']}, {'/test/doc/file1': ['http://ulr/somefile.zip', 'http://another file.zip']}, {'/test1': 'http://newfile.tar'} ] def parse_tasks(task_list): tasks = [] for item in task_list: for dest_path, urls in item.items(): urls = [urls] if isinstance(urls, str) else urls for url in urls: tasks.append((dest_path.strip(), url.strip())) return tasks def download(dest_path, url, semaphore): with semaphore: try: dest = Path(dest_path) if dest.suffix == '': dest.mkdir(parents=True, exist_ok=True) wget.download(url, out=str(dest)) else: dest.parent.mkdir(parents=True, exist_ok=True) wget.download(url, out=str(dest)) print(f"\n下载完成: {url}") except Exception as e: print(f"\n下载失败 {url}: {str(e)}") if __name__ == "__main__": max_concurrent = 10 semaphore = threading.Semaphore(max_concurrent) tasks = parse_tasks(testList) threads = [] for dest, url in tasks: thread = threading.Thread(target=download, args=(dest, url, semaphore)) threads.append(thread) thread.start() # 等待所有线程执行完毕 for thread in threads: thread.join() print("所有下载任务完成")
方案3:异步IO下载(适合超大量任务)
如果需要更高的并发效率,可以用异步IO方案,依赖aiohttp和aiofiles库(需提前安装:pip install aiohttp aiofiles)。
示例代码
import asyncio import aiohttp import aiofiles from pathlib import Path testList = [ {'/test/doc/': ['http://firstfile.tar', 'http://secondfile.tar']}, {'/test/doc/file1': ['http://ulr/somefile.zip', 'http://another file.zip']}, {'/test1': 'http://newfile.tar'} ] def parse_tasks(task_list): tasks = [] for item in task_list: for dest_path, urls in item.items(): urls = [urls] if isinstance(urls, str) else urls for url in urls: tasks.append((dest_path.strip(), url.strip())) return tasks async def async_download(session, dest_path, url): try: dest = Path(dest_path) # 处理目标路径:文件夹则自动生成文件名,文件路径则直接使用 if dest.suffix == '': dest.mkdir(parents=True, exist_ok=True) filename = url.split('/')[-1].split('?')[0] dest_file = dest / filename else: dest_file = dest dest_file.parent.mkdir(parents=True, exist_ok=True) async with session.get(url) as response: if response.status == 200: async with aiofiles.open(dest_file, 'wb') as f: await f.write(await response.read()) print(f"下载完成: {url} -> {dest_file}") else: print(f"下载失败 {url}: HTTP状态码 {response.status}") except Exception as e: print(f"下载失败 {url}: {str(e)}") async def main(max_concurrent=10): tasks = parse_tasks(testList) # 用信号量限制异步并发数 semaphore = asyncio.Semaphore(max_concurrent) async def bounded_download(dest, url): async with semaphore: async with aiohttp.ClientSession() as session: await async_download(session, dest, url) await asyncio.gather(*[bounded_download(dest, url) for dest, url in tasks]) if __name__ == "__main__": asyncio.run(main(max_concurrent=10)) print("所有下载任务完成")
总结
- 对于下载这类IO密集型任务,多线程是最直接高效的选择,GIL不会成为瓶颈
ThreadPoolExecutor代码最简洁,推荐优先使用- 异步IO适合超大量任务场景,性能优势更明显
内容的提问来源于stack exchange,提问作者Sunny
相关产品推荐
相关产品推荐

