如何用Python多线程快速下载1000+个.txt文件并实现词索引
当然可以!用Python并发下载彻底解决瓶颈
你的核心问题是串行下载的IO阻塞——网络请求属于IO密集型任务,Python的并发模型(多线程/异步IO)刚好能完美应对,在4核机器上完全能把总运行时间压到接近索引的耗时水平。
先给你算笔账:串行1000个文件,每个0.25-0.5秒,总耗时是250-500秒;如果用10线程并发,理论上能降到25-50秒;要是网络允许,开到20线程就能进一步压缩到12.5-25秒——这已经比串行快了一个数量级,再加上本身就很快的索引步骤,整体耗时会非常接近你的期望。
下面给你两种最实用的优化方案,附代码示例:
方案1:线程池(入门友好,快速改造)
用concurrent.futures.ThreadPoolExecutor是最容易上手的方式,不需要大幅改动原有代码,直接把下载+索引逻辑包装成任务丢给线程池即可:
import urllib.request from concurrent.futures import ThreadPoolExecutor from collections import defaultdict import time # 初始化索引字典(用defaultdict简化嵌套操作) term_index = defaultdict(lambda: defaultdict(int)) def process_single_url(url): try: # 加个User-Agent避免被目标网站拦截 headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} req = urllib.request.Request(url, headers=headers) with urllib.request.urlopen(req, timeout=10) as response: text = response.read().decode('utf-8', errors='ignore') # 假设从URL提取大学名称(根据你的实际URL格式调整) university_name = url.split('/')[-1].replace('.txt', '') # 分词并更新索引(这里做了小写化,你可以加停用词过滤、标点清理等) for term in text.lower().split(): term_index[term][university_name] += 1 return True except Exception as e: print(f"处理URL失败 {url}: {str(e)}") return False if __name__ == '__main__': start = time.time() # 读取URL列表 with open('urls.txt', 'r') as f: urls = [line.strip() for line in f if line.strip()] # 线程池并发数:4核机器开10-20都没问题(根据网络稳定性调整) with ThreadPoolExecutor(max_workers=15) as executor: executor.map(process_single_url, urls) print(f"总耗时:{time.time() - start:.2f} 秒")
为什么选线程池?
- IO密集型任务下,线程的开销远低于进程,4核机器能轻松承载数十个线程;
- 代码改动极小,原有下载和索引逻辑几乎不用改;
- 自带任务调度和线程管理,不用自己手动维护线程生命周期。
方案2:异步IO(高并发场景更高效)
如果你的网络条件很好,想追求极致速度,可以用aiohttp做异步请求——异步模型没有线程切换的开销,在大量网络请求下效率会更高:
import aiohttp import asyncio from collections import defaultdict import time term_index = defaultdict(lambda: defaultdict(int)) async def process_single_url(session, url): try: headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} async with session.get(url, headers=headers, timeout=10) as response: text = await response.text(encoding='utf-8', errors='ignore') university_name = url.split('/')[-1].replace('.txt', '') for term in text.lower().split(): term_index[term][university_name] += 1 return True except Exception as e: print(f"处理URL失败 {url}: {str(e)}") return False async def main(): start = time.time() with open('urls.txt', 'r') as f: urls = [line.strip() for line in f if line.strip()] # 创建异步Session,复用连接池 async with aiohttp.ClientSession() as session: tasks = [process_single_url(session, url) for url in urls] await asyncio.gather(*tasks) print(f"总耗时:{time.time() - start:.2f} 秒") if __name__ == '__main__': asyncio.run(main())
关键注意事项
- 异常处理:一定要加
try-except捕获网络错误(超时、连接失败、编码错误等),避免单个URL失败导致整个程序崩溃; - 反爬策略:加
User-Agent模拟浏览器请求,必要时可以加随机延迟(比如time.sleep(random.uniform(0.1, 0.3))),避免被目标网站封禁; - 索引优化:你用的嵌套字典没问题,用
collections.defaultdict能让代码更简洁;如果后续数据量更大,可以考虑用pandas或者数据库来存储索引,效率会更高; - 并发数控制:不要盲目开太多线程/异步任务,否则可能会被目标网站限流,或者导致本地网络拥堵——4核机器开10-20线程,异步任务开30-50个是比较稳妥的范围。
最终结论
在你的4核机器上,Python完全能通过并发下载把总运行时间大幅压缩,甚至如果网络条件理想,总耗时能非常接近你提到的1.5秒索引基准。优先推荐线程池方案,上手快、改动小;如果追求极致速度,再尝试异步IO。
内容的提问来源于stack exchange,提问作者BabaSvoloch
相关产品推荐
相关产品推荐

