You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用gevent协程优化词频统计程序却变慢,求问题排查

问题分析与解决方案

你的协程版本运行变慢主要有三个核心原因:

1. 未对标准库打Monkey Patch

Gevent的协程调度依赖于将标准库中的阻塞IO操作(比如文件读写、系统调用)替换为可触发协程切换的异步版本。你当前的代码没有执行补丁,导致文件读写、os.listdir等操作都是阻塞式的,协程根本不会在IO等待时切换,相当于还是串行执行,反而多了协程创建和调度的额外开销。

2. 任务包含大量CPU密集操作

count_words里的正则分割、Counter统计、排序都是CPU密集型任务。Python的GIL(全局解释器锁)会让同一时间只有一个协程执行CPU任务,协程无法利用多核,也不会在CPU计算时切换,这部分任务用协程完全没有优势,反而因为调度拖慢速度。

3. 无限制创建协程

每个文件都创建一个协程,当文件数量较多时,大量协程的调度开销会抵消IO等待的收益,尤其是小文件场景下,IO等待时间远小于协程调度时间。


修复后的代码示例

import os
import re
import time
from collections import Counter
from gevent import monkey, Pool

# 必须先打补丁,替换标准库的阻塞IO操作
monkey.patch_all()

def count_words(content):
    words = [word.strip('"') for word in re.split("[\s-]", content) if re.fullmatch("[a-zA-Z']+", word.strip('"'))]
    word_count = Counter(words)
    return word_count


def process_single_document(folder_name, document_path):
    with open(document_path, mode='r+', encoding='utf8') as f:
        content = f.read()
        word_count = count_words(content)
        content_append = '\n'.join([f'{word} {count}' for word, count in sorted(word_count.items(), key=lambda pair: pair[0])])
        f.write(content_append)
        file_name = f'{folder_name}_{os.path.basename(document_path)}'
        return {file_name: sum(word_count.values())}

def timecost(func):
    def wrapper(*args, **kwargs):
        start = time.time()
        res = func(*args, **kwargs)
        print(f'{func.__name__} 耗时:{time.time() - start}')
        return res
    return wrapper

@timecost
def process_single_folder(folder_path):
    folder_name = os.path.basename(folder_path)
    folder_res = {}
    for file in os.listdir(folder_path):
        file_path = os.path.join(folder_path, file)
        folder_res.update(process_single_document(folder_name, file_path))
    print(f'file counts: {len(folder_res)}')
    return folder_res

@timecost
def process_single_folder_coroutine(folder_path):
    folder_name = os.path.basename(folder_path)
    # 使用Pool限制协程数量,IO密集场景可设为CPU核心数*2或稍大值
    pool = Pool(10)
    tasks = []
    for file in os.listdir(folder_path):
        file_path = os.path.join(folder_path, file)
        tasks.append(pool.spawn(process_single_document, folder_name, file_path))
    pool.join()
    folder_res = {}
    for task in tasks:
        folder_res.update(task.value)
    print(f'file counts: {len(folder_res)}')
    return folder_res

额外优化建议

如果CPU密集部分占比高,建议结合多进程+协程:用多进程处理CPU密集的单词统计,协程处理文件IO,这样既能利用多核,又能优化IO等待时间。比如用multiprocessing.Pool来并行执行count_words函数。

内容的提问来源于stack exchange,提问作者jiligulu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 12:39:51