大Pandas DataFrame词频统计优化:多进程/多线程方案咨询
高效统计Pandas文本列的词频
原方案慢的核心原因
你的代码性能瓶颈在于两点:
- 拼接所有句子为超大字符串再拆分,既浪费内存又完全没必要
text_as_list.count(w)属于O(n²)级操作:对每个唯一单词都要遍历整个单词列表,200万单词级别的数据会导致计算量爆炸
单进程最优方案
直接用Python标准库或Pandas的底层优化实现,性能远超纯Python循环:
方法1:使用collections.Counter
跳过拼接大字符串的步骤,直接迭代所有句子的单词:
from collections import Counter # 生成所有单词的迭代器,传入Counter统计 word_counts = Counter(word for sentence in df["sentence"] for word in sentence.split()) # 转为排序后的字典 sorted_word_counts = dict(sorted(word_counts.items()))
方法2:使用Pandas向量化操作
利用Pandas的字符串处理和堆叠统计,速度同样高效:
# 拆分每个句子为单词,堆叠成单列后统计词频 word_counts_df = df["sentence"].str.split(expand=True).stack().value_counts() # 转为排序后的字典 sorted_word_counts = dict(sorted(word_counts_df.to_dict().items()))
多进程/多线程的选择
这个任务属于CPU密集型:核心工作是字符串拆分和计数计算,Python的GIL(全局解释器锁)会导致多线程无法利用多核CPU,因此多进程更合适——ProcessPoolExecutor可以绕过GIL,让每个进程占用一个CPU核心,并行处理数据块。
基于ProcessPoolExecutor的多进程实现
思路:将DataFrame拆分为多个子块,每个子块单独统计词频,最后合并所有子块的统计结果:
from collections import Counter from concurrent.futures import ProcessPoolExecutor import pandas as pd # 单个数据块的统计函数 def count_chunk(chunk): return Counter(word for sentence in chunk for word in sentence.split()) def parallel_word_counter(df, column_name, num_workers=None): # 根据进程数拆分DataFrame为子块 chunk_size = len(df) // (num_workers or 4) chunks = [chunk for _, chunk in df.groupby(df.index // chunk_size)] # 启动多进程池并行处理 with ProcessPoolExecutor(max_workers=num_workers) as executor: chunk_counters = list(executor.map(count_chunk, [chunk[column_name] for chunk in chunks])) # 合并所有子块的统计结果 total_counter = Counter() for counter in chunk_counters: total_counter.update(counter) # 返回排序后的字典 return dict(sorted(total_counter.items())) # 调用示例 sorted_word_counts = parallel_word_counter(df, "sentence")
注意事项
num_workers建议设置为CPU核心数(可通过os.cpu_count()获取),避免过度创建进程导致资源浪费- 如果需要处理特殊格式文本(如统一大小写、去除标点),可在
count_chunk函数中添加预处理逻辑(比如sentence.lower().strip(',.!?'))
内容的提问来源于stack exchange,提问作者Miguel Garcia
相关产品推荐
相关产品推荐

