You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas绘制按标签分组的特征分布重叠直方图?

按标签在同一子图展示特征分布的实现方法

方案1:用Seaborn快速实现(推荐)

Seaborn的histplot支持通过hue参数直接按标签分组绘制,代码简洁易读:

import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt

# 生成示例数据
x1 = np.random.randn(1000)
x2 = np.random.randn(1000)
y = np.random.randint(0,2,size=1000)
df = pd.DataFrame({'x1':x1, 'x2':x2, 'y': y.astype(str)}) # 将y转为字符串,让图例显示更清晰

# 创建1行2列的子图布局
fig, axes = plt.subplots(1, 2, figsize=(10, 4))

# 绘制x1的分组分布
sns.histplot(data=df, x='x1', hue='y', bins=10, ax=axes[0], kde=False)
axes[0].set_title('x1 按标签y的分布')

# 绘制x2的分组分布
sns.histplot(data=df, x='x2', hue='y', bins=10, ax=axes[1], kde=False)
axes[1].set_title('x2 按标签y的分布')

plt.tight_layout()
plt.show()

这段代码会生成两个子图,每个子图里同时展示标签0和1对应的特征分布,自动生成图例区分不同标签。

方案2:用Matplotlib原生代码实现

如果不想依赖Seaborn,可以手动拆分数据后绘制:

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# 生成示例数据
x1 = np.random.randn(1000)
x2 = np.random.randn(1000)
y = np.random.randint(0,2,size=1000)
df = pd.DataFrame({'x1':x1, 'x2':x2, 'y': y})

# 创建1行2列的子图
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 4))

# 绘制x1的两个标签分布,用alpha设置透明度避免遮挡
ax1.hist(df[df['y']==0]['x1'], bins=10, alpha=0.5, label='y=0')
ax1.hist(df[df['y']==1]['x1'], bins=10, alpha=0.5, label='y=1')
ax1.set_title('x1 分布')
ax1.legend()

# 绘制x2的两个标签分布
ax2.hist(df[df['y']==0]['x2'], bins=10, alpha=0.5, label='y=0')
ax2.hist(df[df['y']==1]['x2'], bins=10, alpha=0.5, label='y=1')
ax2.set_title('x2 分布')
ax2.legend()

plt.tight_layout()
plt.show()

通过布尔索引拆分不同标签的数据,叠加绘制直方图,用alpha参数调整透明度保证两个分布都能看清,最后添加图例区分标签。

内容的提问来源于stack exchange,提问作者Clovis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 18:54:58