You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python WordCloud库时字母"S"被计入词频的原因及解决方案咨询

解决WordCloud统计NLTK《白鲸记》中单独"S"词频的问题

嘿,这个问题我碰到过类似的!咱们一步步来拆解和解决:

问题根源

你遇到的核心问题是NLTK的gutenberg语料库把所有格形式拆成了两个独立的词——比如原本的whale's被拆成了whale和/'s。当你直接把所有词拼接成字符串后,WordCloud的默认分词规则会把/'s里的s(或S)识别成单独的词,自然就会被统计成高频词。而你参考的教程大概率是直接读取完整的文本文件(没有经过NLTK的words()拆分),所以不会出现这个拆分问题。

两种解决方案

方案1:预处理文本,合并拆分的所有格(推荐)

这是最彻底的方法,直接修复语料里的所有格拆分问题,从源头避免错误统计:

import nltk
from wordcloud import WordCloud
import matplotlib.pyplot as plt

# 先下载语料(如果没下载过)
nltk.download('gutenberg')
example_corpus = nltk.corpus.gutenberg.words("melville-moby_dick.txt")

# 处理拆分的所有格形式
processed_words = []
i = 0
while i < len(example_corpus):
    current_word = example_corpus[i]
    # 检查下一个词是否是所有格标记(/'s 或 /'S)
    if i + 1 < len(example_corpus) and example_corpus[i+1] in ("/'s", "/'S"):
        # 合并成正常的所有格形式,比如 whale/'s → whale's
        processed_words.append(f"{current_word}'s")
        i += 2  # 跳过下一个所有格标记词
    else:
        processed_words.append(current_word)
        i += 1

# 拼接成完整文本
novel_as_string = " ".join(processed_words)

# 生成词云
wordcloud = WordCloud().generate(novel_as_string)

plt.imshow(wordcloud, interpolation="bilinear")
plt.axis("off")
plt.show()

方案2:将单独的"S"/"s"加入停用词(快速临时解决)

如果你不想做复杂的预处理,也可以直接把单独的s和S加入WordCloud的停用词列表,让它忽略这些词:

import nltk
from wordcloud import WordCloud
import matplotlib.pyplot as plt

nltk.download('gutenberg')
example_corpus = nltk.corpus.gutenberg.words("melville-moby_dick.txt")
novel_as_string = " ".join(example_corpus)

# 自定义停用词,添加单独的s和S
custom_stopwords = WordCloud().stopwords.union({'s', 'S'})
wordcloud = WordCloud(stopwords=custom_stopwords).generate(novel_as_string)

plt.imshow(wordcloud, interpolation="bilinear")
plt.axis("off")
plt.show()

不过这个方法有局限性:如果文本里真的有需要统计的单独S(比如缩写),也会被过滤掉,所以还是方案1更准确。

内容的提问来源于stack exchange,提问作者Caelan Curry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 21:02:28