鸢尾花数据集柱状图物种类别颜色重叠问题求助
解决鸢尾花数据集柱状图颜色重叠问题
你当前的代码会把每个样本的数值单独画成柱子,同一物种的所有样本柱子完全重叠在一起;同时color=c会循环给每个柱子分配颜色,导致每个物种的柱状条上叠加了三种颜色,这就是你看到的效果。
要实现不同物种用不同颜色区分,需要先对每个物种的特征做统计汇总(比如平均值),让每个物种对应一个柱子,再给每个柱子分配对应颜色。
修改后的代码:
import pandas as pd import matplotlib.pyplot as plt # 假设df是加载好的鸢尾花数据集 df = pd.read_csv('iris.csv') # 按物种分组计算特征均值 species_grouped = df.groupby('Species').mean() c = ['red', 'green', 'orange'] species_labels = species_grouped.index fig, axes = plt.subplots(2, 2, figsize=(8, 8)) # 绘制花萼长度柱状图 axes[0,0].set_title('Sepal Length') axes[0,0].bar(species_labels, species_grouped['SepalLengthCm'], color=c) # 绘制花萼宽度柱状图 axes[0,1].set_title('Sepal Width') axes[0,1].bar(species_labels, species_grouped['SepalWidthCm'], color=c) # 绘制花瓣长度柱状图 axes[1,0].set_title('Petal Length') axes[1,0].bar(species_labels, species_grouped['PetalLengthCm'], color=c) # 绘制花瓣宽度柱状图 axes[1,1].set_title('Petal Width') axes[1,1].bar(species_labels, species_grouped['PetalWidthCm'], color=c) plt.tight_layout() plt.show()
如果想展示每个物种的特征分布(而非单一统计值),箱线图会更合适,示例代码如下:
import pandas as pd import matplotlib.pyplot as plt df = pd.read_csv('iris.csv') fig, axes = plt.subplots(2, 2, figsize=(8, 8)) features = ['SepalLengthCm', 'SepalWidthCm', 'PetalLengthCm', 'PetalWidthCm'] species_unique = df['Species'].unique() c = ['red', 'green', 'orange'] for ax, feature in zip(axes.flat, features): # 为每个物种的特征数据绘制箱线图 box_plots = ax.boxplot( [df[df['Species'] == sp][feature] for sp in species_unique], labels=species_unique, patch_artist=True ) # 给每个箱子分配对应颜色 for patch, color in zip(box_plots['boxes'], c): patch.set_facecolor(color) ax.set_title(feature) plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者Souparno
相关产品推荐
相关产品推荐

