You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python多线程快速下载1000+个.txt文件并实现词索引

当然可以!用Python并发下载彻底解决瓶颈

你的核心问题是串行下载的IO阻塞——网络请求属于IO密集型任务,Python的并发模型(多线程/异步IO)刚好能完美应对,在4核机器上完全能把总运行时间压到接近索引的耗时水平。

先给你算笔账:串行1000个文件,每个0.25-0.5秒,总耗时是250-500秒;如果用10线程并发,理论上能降到25-50秒;要是网络允许,开到20线程就能进一步压缩到12.5-25秒——这已经比串行快了一个数量级,再加上本身就很快的索引步骤,整体耗时会非常接近你的期望。

下面给你两种最实用的优化方案,附代码示例:

方案1:线程池(入门友好,快速改造)

用concurrent.futures.ThreadPoolExecutor是最容易上手的方式,不需要大幅改动原有代码,直接把下载+索引逻辑包装成任务丢给线程池即可:

import urllib.request
from concurrent.futures import ThreadPoolExecutor
from collections import defaultdict
import time

# 初始化索引字典(用defaultdict简化嵌套操作)
term_index = defaultdict(lambda: defaultdict(int))

def process_single_url(url):
    try:
        # 加个User-Agent避免被目标网站拦截
        headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
        req = urllib.request.Request(url, headers=headers)
        
        with urllib.request.urlopen(req, timeout=10) as response:
            text = response.read().decode('utf-8', errors='ignore')
            
            # 假设从URL提取大学名称(根据你的实际URL格式调整)
            university_name = url.split('/')[-1].replace('.txt', '')
            
            # 分词并更新索引(这里做了小写化,你可以加停用词过滤、标点清理等)
            for term in text.lower().split():
                term_index[term][university_name] += 1
        return True
    except Exception as e:
        print(f"处理URL失败 {url}: {str(e)}")
        return False

if __name__ == '__main__':
    start = time.time()
    
    # 读取URL列表
    with open('urls.txt', 'r') as f:
        urls = [line.strip() for line in f if line.strip()]
    
    # 线程池并发数:4核机器开10-20都没问题(根据网络稳定性调整)
    with ThreadPoolExecutor(max_workers=15) as executor:
        executor.map(process_single_url, urls)
    
    print(f"总耗时:{time.time() - start:.2f} 秒")

为什么选线程池?

  • IO密集型任务下,线程的开销远低于进程,4核机器能轻松承载数十个线程;
  • 代码改动极小,原有下载和索引逻辑几乎不用改;
  • 自带任务调度和线程管理,不用自己手动维护线程生命周期。

方案2:异步IO(高并发场景更高效)

如果你的网络条件很好,想追求极致速度,可以用aiohttp做异步请求——异步模型没有线程切换的开销,在大量网络请求下效率会更高:

import aiohttp
import asyncio
from collections import defaultdict
import time

term_index = defaultdict(lambda: defaultdict(int))

async def process_single_url(session, url):
    try:
        headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
        async with session.get(url, headers=headers, timeout=10) as response:
            text = await response.text(encoding='utf-8', errors='ignore')
            
            university_name = url.split('/')[-1].replace('.txt', '')
            for term in text.lower().split():
                term_index[term][university_name] += 1
        return True
    except Exception as e:
        print(f"处理URL失败 {url}: {str(e)}")
        return False

async def main():
    start = time.time()
    with open('urls.txt', 'r') as f:
        urls = [line.strip() for line in f if line.strip()]
    
    # 创建异步Session,复用连接池
    async with aiohttp.ClientSession() as session:
        tasks = [process_single_url(session, url) for url in urls]
        await asyncio.gather(*tasks)
    
    print(f"总耗时:{time.time() - start:.2f} 秒")

if __name__ == '__main__':
    asyncio.run(main())

关键注意事项

  1. 异常处理:一定要加try-except捕获网络错误(超时、连接失败、编码错误等),避免单个URL失败导致整个程序崩溃;
  2. 反爬策略:加User-Agent模拟浏览器请求,必要时可以加随机延迟(比如time.sleep(random.uniform(0.1, 0.3))),避免被目标网站封禁;
  3. 索引优化:你用的嵌套字典没问题,用collections.defaultdict能让代码更简洁;如果后续数据量更大,可以考虑用pandas或者数据库来存储索引,效率会更高;
  4. 并发数控制:不要盲目开太多线程/异步任务,否则可能会被目标网站限流,或者导致本地网络拥堵——4核机器开10-20线程,异步任务开30-50个是比较稳妥的范围。

最终结论

在你的4核机器上,Python完全能通过并发下载把总运行时间大幅压缩,甚至如果网络条件理想,总耗时能非常接近你提到的1.5秒索引基准。优先推荐线程池方案,上手快、改动小;如果追求极致速度,再尝试异步IO。

内容的提问来源于stack exchange,提问作者BabaSvoloch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:33:05