You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大Pandas DataFrame词频统计优化:多进程/多线程方案咨询

高效统计Pandas文本列的词频

原方案慢的核心原因

你的代码性能瓶颈在于两点:

  • 拼接所有句子为超大字符串再拆分,既浪费内存又完全没必要
  • text_as_list.count(w) 属于O(n²)级操作:对每个唯一单词都要遍历整个单词列表,200万单词级别的数据会导致计算量爆炸

单进程最优方案

直接用Python标准库或Pandas的底层优化实现,性能远超纯Python循环:

方法1:使用collections.Counter

跳过拼接大字符串的步骤,直接迭代所有句子的单词:

from collections import Counter

# 生成所有单词的迭代器,传入Counter统计
word_counts = Counter(word for sentence in df["sentence"] for word in sentence.split())
# 转为排序后的字典
sorted_word_counts = dict(sorted(word_counts.items()))

方法2:使用Pandas向量化操作

利用Pandas的字符串处理和堆叠统计,速度同样高效:

# 拆分每个句子为单词,堆叠成单列后统计词频
word_counts_df = df["sentence"].str.split(expand=True).stack().value_counts()
# 转为排序后的字典
sorted_word_counts = dict(sorted(word_counts_df.to_dict().items()))

多进程/多线程的选择

这个任务属于CPU密集型:核心工作是字符串拆分和计数计算,Python的GIL(全局解释器锁)会导致多线程无法利用多核CPU,因此多进程更合适——ProcessPoolExecutor可以绕过GIL,让每个进程占用一个CPU核心,并行处理数据块。

基于ProcessPoolExecutor的多进程实现

思路:将DataFrame拆分为多个子块,每个子块单独统计词频,最后合并所有子块的统计结果:

from collections import Counter
from concurrent.futures import ProcessPoolExecutor
import pandas as pd

# 单个数据块的统计函数
def count_chunk(chunk):
    return Counter(word for sentence in chunk for word in sentence.split())

def parallel_word_counter(df, column_name, num_workers=None):
    # 根据进程数拆分DataFrame为子块
    chunk_size = len(df) // (num_workers or 4)
    chunks = [chunk for _, chunk in df.groupby(df.index // chunk_size)]
    
    # 启动多进程池并行处理
    with ProcessPoolExecutor(max_workers=num_workers) as executor:
        chunk_counters = list(executor.map(count_chunk, [chunk[column_name] for chunk in chunks]))
    
    # 合并所有子块的统计结果
    total_counter = Counter()
    for counter in chunk_counters:
        total_counter.update(counter)
    
    # 返回排序后的字典
    return dict(sorted(total_counter.items()))

# 调用示例
sorted_word_counts = parallel_word_counter(df, "sentence")

注意事项

  • num_workers 建议设置为CPU核心数(可通过os.cpu_count()获取),避免过度创建进程导致资源浪费
  • 如果需要处理特殊格式文本(如统一大小写、去除标点),可在count_chunk函数中添加预处理逻辑(比如sentence.lower().strip(',.!?'))

内容的提问来源于stack exchange,提问作者Miguel Garcia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 13:22:22