You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于字典列表实现多文件多线程下载?含GIL相关疑问

多文件并发下载解决方案

先澄清GIL的误区

Python的GIL确实会限制CPU密集型任务的多线程性能,但下载属于IO密集型任务——在等待网络响应时,线程会主动释放GIL,所以多线程完全适合这类场景,能有效提升并发效率,不用被GIL的说法束缚。

方案1:用ThreadPoolExecutor实现多线程下载(推荐)

concurrent.futures.ThreadPoolExecutor是Python标准库中更简洁的并发工具,不需要手动管理线程和信号量,直接指定最大并发数即可。

步骤说明

  1. 先把原始字典列表解析成扁平的任务元组(目标路径+单个URL),方便批量处理
  2. 定义下载函数,处理路径创建和文件下载逻辑
  3. 用线程池控制最多10个并发任务

示例代码

import wget
from pathlib import Path
from concurrent.futures import ThreadPoolExecutor

# 你的原始任务列表
testList = [
    {'/test/doc/': ['http://firstfile.tar', 'http://secondfile.tar']},
    {'/test/doc/file1': ['http://ulr/somefile.zip', 'http://another file.zip']},
    {'/test1': 'http://newfile.tar'}
]

def parse_tasks(task_list):
    """把嵌套的任务列表转换成扁平的(目标路径, URL)元组列表"""
    tasks = []
    for item in task_list:
        for dest_path, urls in item.items():
            # 兼容单个URL的字符串格式
            urls = [urls] if isinstance(urls, str) else urls
            for url in urls:
                tasks.append((dest_path.strip(), url.strip()))
    return tasks

def download_task(dest_path, url):
    """单个文件的下载逻辑"""
    try:
        dest = Path(dest_path)
        # 如果目标是文件夹,先创建目录,用URL中的文件名保存
        if dest.suffix == '':
            dest.mkdir(parents=True, exist_ok=True)
            wget.download(url, out=str(dest))
        # 如果目标是具体文件路径,直接指定输出
        else:
            dest.parent.mkdir(parents=True, exist_ok=True)
            wget.download(url, out=str(dest))
        print(f"\n下载完成: {url}")
    except Exception as e:
        print(f"\n下载失败 {url}: {str(e)}")

if __name__ == "__main__":
    max_concurrent = 10
    tasks = parse_tasks(testList)
    
    # 启动线程池执行任务
    with ThreadPoolExecutor(max_workers=max_concurrent) as executor:
        executor.map(lambda x: download_task(*x), tasks)
    
    print("所有下载任务完成")

方案2:手动用threading+信号量控制并发

如果需要更精细的线程管理,可以手动创建线程,用Semaphore限制最大并发数。

示例代码

import threading
import wget
from pathlib import Path

testList = [
    {'/test/doc/': ['http://firstfile.tar', 'http://secondfile.tar']},
    {'/test/doc/file1': ['http://ulr/somefile.zip', 'http://another file.zip']},
    {'/test1': 'http://newfile.tar'}
]

def parse_tasks(task_list):
    tasks = []
    for item in task_list:
        for dest_path, urls in item.items():
            urls = [urls] if isinstance(urls, str) else urls
            for url in urls:
                tasks.append((dest_path.strip(), url.strip()))
    return tasks

def download(dest_path, url, semaphore):
    with semaphore:
        try:
            dest = Path(dest_path)
            if dest.suffix == '':
                dest.mkdir(parents=True, exist_ok=True)
                wget.download(url, out=str(dest))
            else:
                dest.parent.mkdir(parents=True, exist_ok=True)
                wget.download(url, out=str(dest))
            print(f"\n下载完成: {url}")
        except Exception as e:
            print(f"\n下载失败 {url}: {str(e)}")

if __name__ == "__main__":
    max_concurrent = 10
    semaphore = threading.Semaphore(max_concurrent)
    tasks = parse_tasks(testList)
    
    threads = []
    for dest, url in tasks:
        thread = threading.Thread(target=download, args=(dest, url, semaphore))
        threads.append(thread)
        thread.start()
    
    # 等待所有线程执行完毕
    for thread in threads:
        thread.join()
    
    print("所有下载任务完成")

方案3:异步IO下载(适合超大量任务)

如果需要更高的并发效率,可以用异步IO方案,依赖aiohttp和aiofiles库(需提前安装:pip install aiohttp aiofiles)。

示例代码

import asyncio
import aiohttp
import aiofiles
from pathlib import Path

testList = [
    {'/test/doc/': ['http://firstfile.tar', 'http://secondfile.tar']},
    {'/test/doc/file1': ['http://ulr/somefile.zip', 'http://another file.zip']},
    {'/test1': 'http://newfile.tar'}
]

def parse_tasks(task_list):
    tasks = []
    for item in task_list:
        for dest_path, urls in item.items():
            urls = [urls] if isinstance(urls, str) else urls
            for url in urls:
                tasks.append((dest_path.strip(), url.strip()))
    return tasks

async def async_download(session, dest_path, url):
    try:
        dest = Path(dest_path)
        # 处理目标路径:文件夹则自动生成文件名,文件路径则直接使用
        if dest.suffix == '':
            dest.mkdir(parents=True, exist_ok=True)
            filename = url.split('/')[-1].split('?')[0]
            dest_file = dest / filename
        else:
            dest_file = dest
            dest_file.parent.mkdir(parents=True, exist_ok=True)
        
        async with session.get(url) as response:
            if response.status == 200:
                async with aiofiles.open(dest_file, 'wb') as f:
                    await f.write(await response.read())
                print(f"下载完成: {url} -> {dest_file}")
            else:
                print(f"下载失败 {url}: HTTP状态码 {response.status}")
    except Exception as e:
        print(f"下载失败 {url}: {str(e)}")

async def main(max_concurrent=10):
    tasks = parse_tasks(testList)
    # 用信号量限制异步并发数
    semaphore = asyncio.Semaphore(max_concurrent)
    
    async def bounded_download(dest, url):
        async with semaphore:
            async with aiohttp.ClientSession() as session:
                await async_download(session, dest, url)
    
    await asyncio.gather(*[bounded_download(dest, url) for dest, url in tasks])

if __name__ == "__main__":
    asyncio.run(main(max_concurrent=10))
    print("所有下载任务完成")

总结

  • 对于下载这类IO密集型任务,多线程是最直接高效的选择,GIL不会成为瓶颈
  • ThreadPoolExecutor代码最简洁,推荐优先使用
  • 异步IO适合超大量任务场景,性能优势更明显

内容的提问来源于stack exchange,提问作者Sunny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 18:20:34