基于NLTK处理chesterton-brown.txt的文本分析技术咨询
问题解决与任务实现指南
先解决你的核心疑问:为什么gutenberg.words('chesterton-brown.txt')只返回6个单词?
大概率是你直接打印结果时,Python自动截断了输出(默认只显示前几项)。你可以用len(brown_words)查看总数量——正常情况下,这个语料库中的《布朗神父》文本单词数是几万级别的。如果数量确实异常,检查NLTK语料库是否下载完整,或改用本地文件加载方式(见下文)。
整体任务实现流程
1. 导入依赖并加载文本
先确保所有需要的NLTK资源都已下载,再正确加载文本:
from nltk.corpus import gutenberg import nltk import string from collections import Counter import matplotlib.pyplot as plt # 首次运行需下载资源 nltk.download('gutenberg') nltk.download('stopwords') nltk.download('punkt') # 方式1:从NLTK Gutenberg语料库加载 brown_words = gutenberg.words('chesterton-brown.txt') # 方式2:加载本地文本文件(如果你的文件是本地存储) # with open('chesterton-brown.txt', 'r', encoding='utf-8') as f: # text = f.read() # brown_words = nltk.word_tokenize(text)
2. 统计文本总单词数量
直接用len()获取原始单词数(包含标点、停用词、重复词):
total_words = len(brown_words) print(f"原始文本总单词数: {total_words}")
3. 找出10个最常用单词并绘制柱状图
统计高频词
用Counter快速统计词频:
word_counts = Counter(brown_words) top10_raw = word_counts.most_common(10) print("原始前10高频词(含标点、停用词):") for word, count in top10_raw: print(f"{word}: {count}")
绘制柱状图
words_raw, counts_raw = zip(*top10_raw) plt.figure(figsize=(10, 6)) plt.bar(words_raw, counts_raw, color='skyblue') plt.title('Top 10 最常用单词(含标点、停用词)') plt.xlabel('单词') plt.ylabel('出现次数') plt.xticks(rotation=45) plt.show()
4. 移除停用词和标点后,再次统计并绘图
文本预处理
先过滤掉停用词和标点:
# 获取英文停用词集合 stop_words = set(nltk.corpus.stopwords.words('english')) # 扩展标点集合,覆盖文本中可能出现的特殊符号 punctuations = set(string.punctuation) punctuations.update({'--', '’', '“', '”', '...'}) # 预处理:转小写、去停用词、去标点 processed_words = [ word.lower() for word in brown_words if word.lower() not in stop_words and word not in punctuations and word.strip() != '' ]
统计处理后的高频词
processed_counts = Counter(processed_words) top10_processed = processed_counts.most_common(10) print("\n移除停用词和标点后的前10高频词:") for word, count in top10_processed: print(f"{word}: {count}")
绘制处理后的柱状图
words_proc, counts_proc = zip(*top10_processed) plt.figure(figsize=(10, 6)) plt.bar(words_proc, counts_proc, color='salmon') plt.title('Top 10 最常用单词(已移除停用词和标点)') plt.xlabel('单词') plt.ylabel('出现次数') plt.xticks(rotation=45) plt.show()
关键函数说明
gutenberg.words(fileid):从NLTK Gutenberg语料库加载指定文本的单词列表nltk.word_tokenize(text):对本地文本进行精准分词nltk.corpus.stopwords.words('english'):获取标准英文停用词列表collections.Counter():高效统计元素出现频率Counter.most_common(n):提取出现次数最多的前n个元素matplotlib.pyplot.bar():生成柱状图可视化词频
内容的提问来源于stack exchange,提问作者Dima
相关产品推荐
相关产品推荐

