如何基于NLTK高效统计2.3M行DataFrame的目标高频词?
高效提取230万行DataFrame高频词的优化方案
原代码的核心低效问题
原代码时间复杂度为O(M*N)(M为行数,N为唯一词数),这是运行极慢的根本原因:
- 对每个唯一词都全量遍历230万行数据,重复计算所有文本的过滤结果
- 逐行
apply处理未利用批量/并行优化 - 词性标注在循环中重复执行,浪费计算资源
优化方案
方案1:单进程批量处理(基础优化)
仅遍历一次数据集,收集所有符合要求的词汇后统一统计词频,时间复杂度降至O(M):
import re from collections import Counter import nltk from nltk.corpus import stopwords # 下载必要的NLTK资源 nltk.download('punkt') nltk.download('averaged_perceptron_tagger') nltk.download('stopwords') # 定义要排除的词汇集合:冠词+英文停用词 exclude_articles = {'a', 'an', 'the'} stop_words = set(stopwords.words('english')) exclude_words = exclude_articles.union(stop_words) def process_single_text(text): # 提取纯单词并转小写 words = re.findall(r'\b\w+\b', str(text).lower()) # 词性标注 tagged_words = nltk.pos_tag(words) # 过滤规则:排除动词(VB开头标签)、数字(CD标签)、冠词及停用词 filtered_words = [ word for word, pos in tagged_words if not pos.startswith('VB') and pos != 'CD' and word not in exclude_words ] return filtered_words # 批量处理所有文本,收集所有符合条件的词汇 all_valid_words = [] for text in df['Comments_Final']: all_valid_words.extend(process_single_text(text)) # 统计词频并取前100 word_frequency = Counter(all_valid_words) top_100_words = word_frequency.most_common(100) # 输出结果 for word, freq in top_100_words: print(f"{word}: {freq}")
方案2:多进程并行处理(进阶优化)
利用多核CPU并行处理文本,进一步提升230万行数据的处理速度:
import re from collections import Counter import nltk from nltk.corpus import stopwords from multiprocessing import Pool, cpu_count # 下载NLTK资源 nltk.download('punkt') nltk.download('averaged_perceptron_tagger') nltk.download('stopwords') exclude_articles = {'a', 'an', 'the'} stop_words = set(stopwords.words('english')) exclude_words = exclude_articles.union(stop_words) def process_single_text(text): words = re.findall(r'\b\w+\b', str(text).lower()) tagged_words = nltk.pos_tag(words) return [ word for word, pos in tagged_words if not pos.startswith('VB') and pos != 'CD' and word not in exclude_words ] if __name__ == '__main__': # 根据CPU核心数创建进程池 pool = Pool(cpu_count()) # 并行处理所有文本 processed_results = pool.map(process_single_text, df['Comments_Final'].tolist()) pool.close() pool.join() # 合并所有结果并统计词频 all_valid_words = [word for sublist in processed_results for word in sublist] word_frequency = Counter(all_valid_words) top_100_words = word_frequency.most_common(100) for word, freq in top_100_words: print(f"{word}: {freq}")
关键优化点说明
- 减少重复计算:仅遍历一次数据集,避免原代码中每个单词都全量扫描的冗余操作
- 高效查询:将排除词转为集合,查询时间从O(n)降至O(1)
- 并行加速:多进程版本利用CPU多核并行处理文本,大幅缩短处理时间
- 简洁统计:使用
collections.Counter直接统计词频,比手动字典计数更高效
内容的提问来源于stack exchange,提问作者user21294168
相关产品推荐
相关产品推荐

