You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

不使用NLTK,如何过滤文本文件中的停用词并统计词频?

实现排除停用词的高效词频统计

核心优化思路

要高效排除停用词,关键是把停用词存储在**集合(set)**中——集合的成员查找操作时间复杂度是O(1),远快于列表的O(n),处理大文本时效率提升明显。同时用with语句管理文件,避免手动关闭文件的资源泄漏问题。

完整实现代码

import re

# 加载停用词到集合(转小写,确保匹配一致)
stop_words = set()
with open('stopwords.txt', 'r') as stop_file:
    for line in stop_file:
        word = line.strip().lower()
        if word:  # 跳过空行
            stop_words.add(word)

word_count = {}
# 处理目标文本文件
with open('documents.txt', 'r') as input_file:
    for line in input_file:
        # 提取所有单词
        words = re.findall(r'\w+', line)
        for word in words:
            word_lower = word.lower()
            # 跳过停用词
            if word_lower in stop_words:
                continue
            # 统计词频(用get方法简化逻辑)
            word_count[word_lower] = word_count.get(word_lower, 0) + 1

# 按单词排序输出
for word in sorted(word_count.keys()):
    print(word, word_count[word])

关键细节说明

  • 停用词集合化:将停用词存入集合,避免每次检查都遍历整个停用词列表,大幅提升查找效率。
  • 大小写统一:把停用词和目标单词都转成小写,确保匹配时不会因大小写差异漏判(比如"The"和"the"视为同一个词)。
  • 简化词频统计:用dict.get(word_lower, 0)替代原有的if-else判断,代码更简洁。
  • 文件安全管理:with语句会在代码块结束后自动关闭文件,无需手动调用close(),避免资源泄漏。
  • 跳过空行:读取停用词时跳过空行,避免集合中加入无效的空字符串。

内容的提问来源于stack exchange,提问作者user14452102

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:05:22