You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于NLTK高效统计2.3M行DataFrame的目标高频词?

高效提取230万行DataFrame高频词的优化方案

原代码的核心低效问题

原代码时间复杂度为O(M*N)(M为行数,N为唯一词数),这是运行极慢的根本原因:

  • 对每个唯一词都全量遍历230万行数据,重复计算所有文本的过滤结果
  • 逐行apply处理未利用批量/并行优化
  • 词性标注在循环中重复执行,浪费计算资源

优化方案

方案1:单进程批量处理(基础优化)

仅遍历一次数据集,收集所有符合要求的词汇后统一统计词频,时间复杂度降至O(M):

import re
from collections import Counter
import nltk
from nltk.corpus import stopwords

# 下载必要的NLTK资源
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')
nltk.download('stopwords')

# 定义要排除的词汇集合:冠词+英文停用词
exclude_articles = {'a', 'an', 'the'}
stop_words = set(stopwords.words('english'))
exclude_words = exclude_articles.union(stop_words)

def process_single_text(text):
    # 提取纯单词并转小写
    words = re.findall(r'\b\w+\b', str(text).lower())
    # 词性标注
    tagged_words = nltk.pos_tag(words)
    # 过滤规则:排除动词(VB开头标签)、数字(CD标签)、冠词及停用词
    filtered_words = [
        word for word, pos in tagged_words
        if not pos.startswith('VB') 
        and pos != 'CD' 
        and word not in exclude_words
    ]
    return filtered_words

# 批量处理所有文本,收集所有符合条件的词汇
all_valid_words = []
for text in df['Comments_Final']:
    all_valid_words.extend(process_single_text(text))

# 统计词频并取前100
word_frequency = Counter(all_valid_words)
top_100_words = word_frequency.most_common(100)

# 输出结果
for word, freq in top_100_words:
    print(f"{word}: {freq}")

方案2:多进程并行处理(进阶优化)

利用多核CPU并行处理文本,进一步提升230万行数据的处理速度:

import re
from collections import Counter
import nltk
from nltk.corpus import stopwords
from multiprocessing import Pool, cpu_count

# 下载NLTK资源
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')
nltk.download('stopwords')

exclude_articles = {'a', 'an', 'the'}
stop_words = set(stopwords.words('english'))
exclude_words = exclude_articles.union(stop_words)

def process_single_text(text):
    words = re.findall(r'\b\w+\b', str(text).lower())
    tagged_words = nltk.pos_tag(words)
    return [
        word for word, pos in tagged_words
        if not pos.startswith('VB') 
        and pos != 'CD' 
        and word not in exclude_words
    ]

if __name__ == '__main__':
    # 根据CPU核心数创建进程池
    pool = Pool(cpu_count())
    # 并行处理所有文本
    processed_results = pool.map(process_single_text, df['Comments_Final'].tolist())
    pool.close()
    pool.join()

    # 合并所有结果并统计词频
    all_valid_words = [word for sublist in processed_results for word in sublist]
    word_frequency = Counter(all_valid_words)
    top_100_words = word_frequency.most_common(100)

    for word, freq in top_100_words:
        print(f"{word}: {freq}")

关键优化点说明

  • 减少重复计算:仅遍历一次数据集,避免原代码中每个单词都全量扫描的冗余操作
  • 高效查询:将排除词转为集合,查询时间从O(n)降至O(1)
  • 并行加速:多进程版本利用CPU多核并行处理文本,大幅缩短处理时间
  • 简洁统计:使用collections.Counter直接统计词频,比手动字典计数更高效

内容的提问来源于stack exchange,提问作者user21294168

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 15:23:12