You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于文本数据集绘制按词频降序排列的词频分布统计图

词频统计及降序可视化实现方案

以下基于Python实现,可直接运行适配你的需求:

1. 所需依赖库

  • pandas:用于读取处理数据集
  • collections.Counter:用于快速统计词频
  • matplotlib:用于绘制可视化图表
    未安装依赖可执行命令:pip install pandas matplotlib

2. 完整实现代码

import pandas as pd
from collections import Counter
import matplotlib.pyplot as plt

# ---------------------- 1. 加载数据集 ----------------------
# 此处为示例数据集构造,实际使用时替换为你的数据集读取代码即可,比如pd.read_csv('你的数据集文件路径.csv')
df = pd.DataFrame({
    'Text': [
        'This is a long string of words',
        'words have many types',
        'each type represents one thing',
        'thing are different',
        'where are these words'
    ]
})

# ---------------------- 2. 文本预处理与词频统计 ----------------------
all_words = []
for text in df['Text']:
    # 统一转小写后按空格拆分单词,可按需新增去除标点、过滤停用词等逻辑
    words = text.lower().split()
    all_words.extend(words)

# 统计词频并按频次降序排序
word_count = Counter(all_words)
sorted_word_count = sorted(word_count.items(), key=lambda x: x[1], reverse=True)

# 拆分出绘图用的单词列表和对应频次列表
words = [item[0] for item in sorted_word_count]
counts = [item[1] for item in sorted_word_count]

# ---------------------- 3. 绘制词频分布图 ----------------------
plt.figure(figsize=(12, 6))
plt.bar(words, counts, color='#3498db')

# 在柱子上方标注具体频次数值
for x, y in enumerate(counts):
    plt.text(x, y + 0.05, str(y), ha='center', fontsize=10)

plt.xlabel('单词', fontsize=12)
plt.ylabel('出现频次', fontsize=12)
plt.title('词频分布统计图(按降序排列)', fontsize=14)
# 单词名称倾斜45度避免重叠
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()

3. 可选优化项

  • 去除文本标点:预处理阶段增加import re,用text = re.sub(r'[^\w\s]', '', text)清理文本后再分词
  • 过滤停用词:可导入nltk停用词库,过滤掉无实际意义的虚词(比如is、of、these等)

内容的提问来源于stack exchange,提问作者Kath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 21:06:03