如何基于文本数据集绘制按词频降序排列的词频分布统计图
词频统计及降序可视化实现方案
以下基于Python实现,可直接运行适配你的需求:
1. 所需依赖库
- pandas:用于读取处理数据集
- collections.Counter:用于快速统计词频
- matplotlib:用于绘制可视化图表
未安装依赖可执行命令:pip install pandas matplotlib
2. 完整实现代码
import pandas as pd from collections import Counter import matplotlib.pyplot as plt # ---------------------- 1. 加载数据集 ---------------------- # 此处为示例数据集构造,实际使用时替换为你的数据集读取代码即可,比如pd.read_csv('你的数据集文件路径.csv') df = pd.DataFrame({ 'Text': [ 'This is a long string of words', 'words have many types', 'each type represents one thing', 'thing are different', 'where are these words' ] }) # ---------------------- 2. 文本预处理与词频统计 ---------------------- all_words = [] for text in df['Text']: # 统一转小写后按空格拆分单词,可按需新增去除标点、过滤停用词等逻辑 words = text.lower().split() all_words.extend(words) # 统计词频并按频次降序排序 word_count = Counter(all_words) sorted_word_count = sorted(word_count.items(), key=lambda x: x[1], reverse=True) # 拆分出绘图用的单词列表和对应频次列表 words = [item[0] for item in sorted_word_count] counts = [item[1] for item in sorted_word_count] # ---------------------- 3. 绘制词频分布图 ---------------------- plt.figure(figsize=(12, 6)) plt.bar(words, counts, color='#3498db') # 在柱子上方标注具体频次数值 for x, y in enumerate(counts): plt.text(x, y + 0.05, str(y), ha='center', fontsize=10) plt.xlabel('单词', fontsize=12) plt.ylabel('出现频次', fontsize=12) plt.title('词频分布统计图(按降序排列)', fontsize=14) # 单词名称倾斜45度避免重叠 plt.xticks(rotation=45) plt.tight_layout() plt.show()
3. 可选优化项
- 去除文本标点:预处理阶段增加
import re,用text = re.sub(r'[^\w\s]', '', text)清理文本后再分词 - 过滤停用词:可导入nltk停用词库,过滤掉无实际意义的虚词(比如is、of、these等)
内容的提问来源于stack exchange,提问作者Kath
相关产品推荐
相关产品推荐

