You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于NLTK处理chesterton-brown.txt的文本分析技术咨询

问题解决与任务实现指南

先解决你的核心疑问:为什么gutenberg.words('chesterton-brown.txt')只返回6个单词?

大概率是你直接打印结果时,Python自动截断了输出(默认只显示前几项)。你可以用len(brown_words)查看总数量——正常情况下,这个语料库中的《布朗神父》文本单词数是几万级别的。如果数量确实异常,检查NLTK语料库是否下载完整,或改用本地文件加载方式(见下文)。


整体任务实现流程

1. 导入依赖并加载文本

先确保所有需要的NLTK资源都已下载,再正确加载文本:

from nltk.corpus import gutenberg
import nltk
import string
from collections import Counter
import matplotlib.pyplot as plt

# 首次运行需下载资源
nltk.download('gutenberg')
nltk.download('stopwords')
nltk.download('punkt')

# 方式1:从NLTK Gutenberg语料库加载
brown_words = gutenberg.words('chesterton-brown.txt')

# 方式2:加载本地文本文件(如果你的文件是本地存储)
# with open('chesterton-brown.txt', 'r', encoding='utf-8') as f:
#     text = f.read()
# brown_words = nltk.word_tokenize(text)

2. 统计文本总单词数量

直接用len()获取原始单词数(包含标点、停用词、重复词):

total_words = len(brown_words)
print(f"原始文本总单词数: {total_words}")

3. 找出10个最常用单词并绘制柱状图

统计高频词

用Counter快速统计词频:

word_counts = Counter(brown_words)
top10_raw = word_counts.most_common(10)

print("原始前10高频词(含标点、停用词):")
for word, count in top10_raw:
    print(f"{word}: {count}")

绘制柱状图

words_raw, counts_raw = zip(*top10_raw)

plt.figure(figsize=(10, 6))
plt.bar(words_raw, counts_raw, color='skyblue')
plt.title('Top 10 最常用单词(含标点、停用词)')
plt.xlabel('单词')
plt.ylabel('出现次数')
plt.xticks(rotation=45)
plt.show()

4. 移除停用词和标点后,再次统计并绘图

文本预处理

先过滤掉停用词和标点:

# 获取英文停用词集合
stop_words = set(nltk.corpus.stopwords.words('english'))
# 扩展标点集合,覆盖文本中可能出现的特殊符号
punctuations = set(string.punctuation)
punctuations.update({'--', '’', '“', '”', '...'})

# 预处理:转小写、去停用词、去标点
processed_words = [
    word.lower() for word in brown_words
    if word.lower() not in stop_words
    and word not in punctuations
    and word.strip() != ''
]

统计处理后的高频词

processed_counts = Counter(processed_words)
top10_processed = processed_counts.most_common(10)

print("\n移除停用词和标点后的前10高频词:")
for word, count in top10_processed:
    print(f"{word}: {count}")

绘制处理后的柱状图

words_proc, counts_proc = zip(*top10_processed)

plt.figure(figsize=(10, 6))
plt.bar(words_proc, counts_proc, color='salmon')
plt.title('Top 10 最常用单词(已移除停用词和标点)')
plt.xlabel('单词')
plt.ylabel('出现次数')
plt.xticks(rotation=45)
plt.show()

关键函数说明

  • gutenberg.words(fileid):从NLTK Gutenberg语料库加载指定文本的单词列表
  • nltk.word_tokenize(text):对本地文本进行精准分词
  • nltk.corpus.stopwords.words('english'):获取标准英文停用词列表
  • collections.Counter():高效统计元素出现频率
  • Counter.most_common(n):提取出现次数最多的前n个元素
  • matplotlib.pyplot.bar():生成柱状图可视化词频

内容的提问来源于stack exchange,提问作者Dima

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 13:10:51