使用gevent协程优化词频统计程序却变慢,求问题排查
问题分析与解决方案
你的协程版本运行变慢主要有三个核心原因:
1. 未对标准库打Monkey Patch
Gevent的协程调度依赖于将标准库中的阻塞IO操作(比如文件读写、系统调用)替换为可触发协程切换的异步版本。你当前的代码没有执行补丁,导致文件读写、os.listdir等操作都是阻塞式的,协程根本不会在IO等待时切换,相当于还是串行执行,反而多了协程创建和调度的额外开销。
2. 任务包含大量CPU密集操作
count_words里的正则分割、Counter统计、排序都是CPU密集型任务。Python的GIL(全局解释器锁)会让同一时间只有一个协程执行CPU任务,协程无法利用多核,也不会在CPU计算时切换,这部分任务用协程完全没有优势,反而因为调度拖慢速度。
3. 无限制创建协程
每个文件都创建一个协程,当文件数量较多时,大量协程的调度开销会抵消IO等待的收益,尤其是小文件场景下,IO等待时间远小于协程调度时间。
修复后的代码示例
import os import re import time from collections import Counter from gevent import monkey, Pool # 必须先打补丁,替换标准库的阻塞IO操作 monkey.patch_all() def count_words(content): words = [word.strip('"') for word in re.split("[\s-]", content) if re.fullmatch("[a-zA-Z']+", word.strip('"'))] word_count = Counter(words) return word_count def process_single_document(folder_name, document_path): with open(document_path, mode='r+', encoding='utf8') as f: content = f.read() word_count = count_words(content) content_append = '\n'.join([f'{word} {count}' for word, count in sorted(word_count.items(), key=lambda pair: pair[0])]) f.write(content_append) file_name = f'{folder_name}_{os.path.basename(document_path)}' return {file_name: sum(word_count.values())} def timecost(func): def wrapper(*args, **kwargs): start = time.time() res = func(*args, **kwargs) print(f'{func.__name__} 耗时:{time.time() - start}') return res return wrapper @timecost def process_single_folder(folder_path): folder_name = os.path.basename(folder_path) folder_res = {} for file in os.listdir(folder_path): file_path = os.path.join(folder_path, file) folder_res.update(process_single_document(folder_name, file_path)) print(f'file counts: {len(folder_res)}') return folder_res @timecost def process_single_folder_coroutine(folder_path): folder_name = os.path.basename(folder_path) # 使用Pool限制协程数量,IO密集场景可设为CPU核心数*2或稍大值 pool = Pool(10) tasks = [] for file in os.listdir(folder_path): file_path = os.path.join(folder_path, file) tasks.append(pool.spawn(process_single_document, folder_name, file_path)) pool.join() folder_res = {} for task in tasks: folder_res.update(task.value) print(f'file counts: {len(folder_res)}') return folder_res
额外优化建议
如果CPU密集部分占比高,建议结合多进程+协程:用多进程处理CPU密集的单词统计,协程处理文件IO,这样既能利用多核,又能优化IO等待时间。比如用multiprocessing.Pool来并行执行count_words函数。
内容的提问来源于stack exchange,提问作者jiligulu
相关产品推荐
相关产品推荐

