Jupyter开发Word Cloud项目无报错但词云无法显示问题
问题现象
- Coursera平台Python速成课程最终项目词云生成代码运行无报错,但始终无法展示词云
- 开发基于Jupyter环境,需预先上传待处理文本文件;重启Jupyter内核后已上传文件不会持久化存储,必须重新执行文件上传代码、重新选择上传目标文本才能继续后续操作
核心问题原因
- 致命缩进错误:词云渲染展示的代码块被错误缩进至
calculate_frequencies函数内部,且位于return cloud.to_array()语句之后。Python函数执行到return就会终止,这部分渲染代码属于永远不会执行的死代码,自然不会弹出词云。 - Jupyter执行顺序问题:重启内核后所有内存变量、已加载的依赖、已上传的临时文件都会被清空,如果没有按单元格从上到下的顺序重新执行,会出现依赖未导入、
file_contents变量不存在/内容为空的情况,也会导致无输出。 - 词频统计逻辑bug:词频统计时存入字典的key是小写格式的单词,但判断单词是否已在字典中时混用了原始格式单词和小写单词做匹配,会导致同个不同大小写的单词被重复记为新单词,词频计数全部为1,影响词云生成准确性。
修复方案
- 先做Jupyter环境前置校验
- 重启内核后,首先执行依赖导入单元格,确保必要依赖已正确加载:
import wordcloud from matplotlib import pyplot as plt- 重新执行文件上传单元格,选中目标文本文件完成上传,执行
print(len(file_contents))校验,若输出值大于0说明文件内容读取正常。
- 修正代码缩进与逻辑bug,完整可用代码如下:
def calculate_frequencies(file_contents): # 预置标点和无意义停用词 punctuations = '''!()-[]{};:'"\,<>./?@#$%^&*_~''' uninteresting_words = ["the", "a", "to", "if", "is", "it", "of", "and", "or", "an", "as", "i", "me", "my", \ "we", "our", "ours", "you", "your", "yours", "he", "she", "him", "his", "her", "hers", "its", "they", "them", \ "their", "what", "which", "who", "whom", "this", "that", "am", "are", "was", "were", "be", "been", "being", \ "have", "has", "had", "do", "does", "did", "but", "at", "by", "with", "from", "here", "when", "where", "how", \ "all", "any", "both", "each", "few", "more", "some", "such", "no", "nor", "too", "very", "can", "will", "just"] # 清理标点 for i in punctuations: file_contents = file_contents.replace(i, '') word_list = file_contents.split() frequency_dict = {} for word in word_list: # 统一转小写后判断,避免大小写导致的计数错误 word_lower = word.lower() if word_lower in uninteresting_words or not word.isalpha(): continue frequency_dict[word_lower] = frequency_dict.get(word_lower, 0) + 1 # 生成词云 cloud = wordcloud.WordCloud() cloud.generate_from_frequencies(frequency_dict) return cloud.to_array() # 词云展示代码必须顶格写,放在函数定义外面,不能缩进在函数内 myimage = calculate_frequencies(file_contents) plt.imshow(myimage, interpolation = 'nearest') plt.axis('off') plt.show()
- 按顺序执行依赖导入、文件上传、函数定义、词云渲染单元格,即可正常显示词云。
注:Jupyter中代码执行顺序严格依赖单元格运行顺序,重启内核后必须从最顶部的单元格开始逐一运行,不要直接跳转到最后一个单元格执行,否则会出现变量不存在、依赖未加载的问题。
内容的提问来源于stack exchange,提问作者DYIII
相关产品推荐
相关产品推荐

