You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按目标变量绘制数据框所有特征的计数(泰坦尼克数据集)

问题描述

我想按目标变量(Survived)绘制数据框中所有特征的计数,完成描述性分析。请问最简单快捷的实现方法是什么?有没有能同时处理数值型和分类数据的方案?我正在用Kaggle的泰坦尼克数据集做实验。

我试了两段代码都有问题:

  1. 这段代码无法显示计数:
sns.pairplot(data=df, diag_kws={'element': 'step', 'histtype': 'step'}, kind='hist', x_vars=['Pclass', 'Sex','Age','SibSp','Parch','Fare','Embarked'], y_vars = ['Survived'])
  1. 这段代码里Age特征的可视化效果混乱,还没法正确设置标签:
# Define the target variable
target_variable = df['Survived']

# Get a list of all other variables
variable_list = ['Pclass', 'Sex','Age','SibSp','Parch','Fare','Embarked']
for i in variable_list:
    m = sns.catplot( x = i, data = df[variable_list], kind = "count", legend = True )
    # Adding Labels to the bars
    ax = m.facet_axis(0,0)
    for p in ax.patches:
        ax.text(p.get_x() - 0.01, 
            p.get_height() * 1.02, 
           '{0:.1f}K'.format((p.get_height()/1000)),   #Used to format it K representation
            color='black', 
            rotation='horizontal', 
            size='large')

解决方案

一、通用方案:自动区分变量类型,按目标分组可视化

结合seaborn和matplotlib,先自动区分分类/数值特征,分别用柱状图、直方图展示,同时按Survived分组并添加计数标签,完美适配泰坦尼克数据集:

import seaborn as sns
import matplotlib.pyplot as plt

# 加载泰坦尼克数据集(本地已有数据可跳过)
df = sns.load_dataset('titanic')

target = 'Survived'
features = ['Pclass', 'Sex','Age','SibSp','Parch','Fare','Embarked']

# 手动区分变量类型(泰坦尼克数据集的特征特性适配)
cat_features = ['Pclass', 'Sex', 'SibSp', 'Parch', 'Embarked']
num_features = ['Age', 'Fare']

# 设置画布
fig, axes = plt.subplots(nrows=len(cat_features)+len(num_features), ncols=1, figsize=(8, 4*(len(cat_features)+len(num_features))))
axes = axes.flatten()

# 处理分类特征:按Survived分组的计数柱状图
for idx, feat in enumerate(cat_features):
    ax = axes[idx]
    sns.countplot(data=df, x=feat, hue=target, ax=ax)
    # 添加计数标签
    for p in ax.patches:
        height = p.get_height()
        if height > 0:
            ax.text(p.get_x() + p.get_width()/2., height,
                    f'{int(height)}',
                    ha='center', va='bottom')
    ax.set_title(f'{feat} 按 {target} 分组计数')

# 处理数值特征:按Survived分组的直方图(带计数)
for idx, feat in enumerate(num_features, start=len(cat_features)):
    ax = axes[idx]
    sns.histplot(data=df, x=feat, hue=target, kde=False, ax=ax, bins=15)
    # 添加计数标签
    for patch in ax.patches:
        height = patch.get_height()
        if height > 0:
            ax.text(patch.get_x() + patch.get_width()/2., height,
                    f'{int(height)}',
                    ha='center', va='bottom')
    ax.set_title(f'{feat} 按 {target} 分组分布计数')

plt.tight_layout()
plt.show()

二、你原有代码的问题修正

  1. 第一段pairplot代码问题:pairplot的设计目的是展示多变量两两关系,不是按目标变量分组的单特征计数,用它做这个需求本身就不合适。
  2. 第二段catplot代码问题:
    • data = df[variable_list]只传入了特征列,没包含Survived,无法按目标分组
    • Age是数值型,用catplot的count类型会把每个唯一值当成单独类别,导致柱子过度密集
    • 计数标签用{0:.1f}K完全没必要,泰坦尼克数据集总样本仅几百,直接显示整数即可

三、快速探索简化方案

如果只追求快速出图、不需要精细美化,可以用pandas原生plot,代码更简洁:

for feat in features:
    # 分类特征/低基数数值特征用分组柱状图
    if df[feat].nunique() < 10 or df[feat].dtype in ['object', 'category']:
        df.groupby(target)[feat].value_counts().unstack().plot(kind='bar', figsize=(8,4), title=f'{feat} by {target}')
    # 高基数数值特征用重叠直方图
    else:
        df.groupby(target)[feat].plot(kind='hist', bins=15, alpha=0.5, legend=True, figsize=(8,4), title=f'{feat} by {target}')
    plt.show()

内容的提问来源于stack exchange,提问作者AbdullahQ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 23:40:25