如何用Pandas绘制Cluster列的分组计数彩色柱状图?
实现DataFrame中Cluster列的差异化颜色柱状图(展示出现次数)
先修正示例数据的问题
原数据存在两个重复的col2键(后一个会覆盖前一个),且Cluster列有拼写不一致的情况("Cluster 3"和"Cluster3"会被识别为不同类别),先做预处理:
import pandas as pd import numpy as np import matplotlib.pyplot as plt # 修正重复列名,整理数据 data = { 'col1': ['Agree', 'Disagree', 'Agree', 'Agree', 'Agree', 'Disagree'], 'col2': ['Agree', 'Disagree', 'Agree', 'Agree', np.nan , 'Disagree'], 'col3': ['Agree', 'Agree', 'Agree', 'Agree', 'Disagree', np.nan], 'Cluster': ['Cluster 1', 'Cluster 2', 'Cluster 2', 'Cluster 1', 'Cluster 3', 'Cluster3'] } df = pd.DataFrame(data) # 统一Cluster命名(可选,若需要合并拼写差异的类别) df['Cluster'] = df['Cluster'].str.replace(' ', '')
步骤1:统计各Cluster的出现次数
用value_counts()快速统计频次:
cluster_counts = df['Cluster'].value_counts()
步骤2:绘制带差异化颜色的柱状图
方法1:Matplotlib 手动指定颜色
可以用colormap生成对应数量的不同颜色:
# 生成与Cluster数量匹配的颜色 colors = plt.cm.tab10(np.arange(len(cluster_counts))) # 绘制柱状图 plt.bar(cluster_counts.index, cluster_counts.values, color=colors) # 添加标注 plt.xlabel('Cluster') plt.ylabel('出现次数') plt.title('各Cluster出现次数统计') plt.show()
方法2:Seaborn 自动分配颜色
Seaborn能更简洁地实现差异化配色:
import seaborn as sns sns.barplot(x=cluster_counts.index, y=cluster_counts.values, palette='tab10') plt.xlabel('Cluster') plt.ylabel('出现次数') plt.title('各Cluster出现次数统计') plt.show()
内容的提问来源于stack exchange,提问作者ar_mm18
相关产品推荐
相关产品推荐

