使用Python WordCloud库时字母"S"被计入词频的原因及解决方案咨询
解决WordCloud统计NLTK《白鲸记》中单独"S"词频的问题
嘿,这个问题我碰到过类似的!咱们一步步来拆解和解决:
问题根源
你遇到的核心问题是NLTK的gutenberg语料库把所有格形式拆成了两个独立的词——比如原本的whale's被拆成了whale和/'s。当你直接把所有词拼接成字符串后,WordCloud的默认分词规则会把/'s里的s(或S)识别成单独的词,自然就会被统计成高频词。而你参考的教程大概率是直接读取完整的文本文件(没有经过NLTK的words()拆分),所以不会出现这个拆分问题。
两种解决方案
方案1:预处理文本,合并拆分的所有格(推荐)
这是最彻底的方法,直接修复语料里的所有格拆分问题,从源头避免错误统计:
import nltk from wordcloud import WordCloud import matplotlib.pyplot as plt # 先下载语料(如果没下载过) nltk.download('gutenberg') example_corpus = nltk.corpus.gutenberg.words("melville-moby_dick.txt") # 处理拆分的所有格形式 processed_words = [] i = 0 while i < len(example_corpus): current_word = example_corpus[i] # 检查下一个词是否是所有格标记(/'s 或 /'S) if i + 1 < len(example_corpus) and example_corpus[i+1] in ("/'s", "/'S"): # 合并成正常的所有格形式,比如 whale/'s → whale's processed_words.append(f"{current_word}'s") i += 2 # 跳过下一个所有格标记词 else: processed_words.append(current_word) i += 1 # 拼接成完整文本 novel_as_string = " ".join(processed_words) # 生成词云 wordcloud = WordCloud().generate(novel_as_string) plt.imshow(wordcloud, interpolation="bilinear") plt.axis("off") plt.show()
方案2:将单独的"S"/"s"加入停用词(快速临时解决)
如果你不想做复杂的预处理,也可以直接把单独的s和S加入WordCloud的停用词列表,让它忽略这些词:
import nltk from wordcloud import WordCloud import matplotlib.pyplot as plt nltk.download('gutenberg') example_corpus = nltk.corpus.gutenberg.words("melville-moby_dick.txt") novel_as_string = " ".join(example_corpus) # 自定义停用词,添加单独的s和S custom_stopwords = WordCloud().stopwords.union({'s', 'S'}) wordcloud = WordCloud(stopwords=custom_stopwords).generate(novel_as_string) plt.imshow(wordcloud, interpolation="bilinear") plt.axis("off") plt.show()
不过这个方法有局限性:如果文本里真的有需要统计的单独S(比如缩写),也会被过滤掉,所以还是方案1更准确。
内容的提问来源于stack exchange,提问作者Caelan Curry
相关产品推荐
相关产品推荐

