如何按目标变量绘制数据框所有特征的计数(泰坦尼克数据集)
问题描述
我想按目标变量(Survived)绘制数据框中所有特征的计数,完成描述性分析。请问最简单快捷的实现方法是什么?有没有能同时处理数值型和分类数据的方案?我正在用Kaggle的泰坦尼克数据集做实验。
我试了两段代码都有问题:
- 这段代码无法显示计数:
sns.pairplot(data=df, diag_kws={'element': 'step', 'histtype': 'step'}, kind='hist', x_vars=['Pclass', 'Sex','Age','SibSp','Parch','Fare','Embarked'], y_vars = ['Survived'])
- 这段代码里Age特征的可视化效果混乱,还没法正确设置标签:
# Define the target variable target_variable = df['Survived'] # Get a list of all other variables variable_list = ['Pclass', 'Sex','Age','SibSp','Parch','Fare','Embarked'] for i in variable_list: m = sns.catplot( x = i, data = df[variable_list], kind = "count", legend = True ) # Adding Labels to the bars ax = m.facet_axis(0,0) for p in ax.patches: ax.text(p.get_x() - 0.01, p.get_height() * 1.02, '{0:.1f}K'.format((p.get_height()/1000)), #Used to format it K representation color='black', rotation='horizontal', size='large')
解决方案
一、通用方案:自动区分变量类型,按目标分组可视化
结合seaborn和matplotlib,先自动区分分类/数值特征,分别用柱状图、直方图展示,同时按Survived分组并添加计数标签,完美适配泰坦尼克数据集:
import seaborn as sns import matplotlib.pyplot as plt # 加载泰坦尼克数据集(本地已有数据可跳过) df = sns.load_dataset('titanic') target = 'Survived' features = ['Pclass', 'Sex','Age','SibSp','Parch','Fare','Embarked'] # 手动区分变量类型(泰坦尼克数据集的特征特性适配) cat_features = ['Pclass', 'Sex', 'SibSp', 'Parch', 'Embarked'] num_features = ['Age', 'Fare'] # 设置画布 fig, axes = plt.subplots(nrows=len(cat_features)+len(num_features), ncols=1, figsize=(8, 4*(len(cat_features)+len(num_features)))) axes = axes.flatten() # 处理分类特征:按Survived分组的计数柱状图 for idx, feat in enumerate(cat_features): ax = axes[idx] sns.countplot(data=df, x=feat, hue=target, ax=ax) # 添加计数标签 for p in ax.patches: height = p.get_height() if height > 0: ax.text(p.get_x() + p.get_width()/2., height, f'{int(height)}', ha='center', va='bottom') ax.set_title(f'{feat} 按 {target} 分组计数') # 处理数值特征:按Survived分组的直方图(带计数) for idx, feat in enumerate(num_features, start=len(cat_features)): ax = axes[idx] sns.histplot(data=df, x=feat, hue=target, kde=False, ax=ax, bins=15) # 添加计数标签 for patch in ax.patches: height = patch.get_height() if height > 0: ax.text(patch.get_x() + patch.get_width()/2., height, f'{int(height)}', ha='center', va='bottom') ax.set_title(f'{feat} 按 {target} 分组分布计数') plt.tight_layout() plt.show()
二、你原有代码的问题修正
- 第一段pairplot代码问题:
pairplot的设计目的是展示多变量两两关系,不是按目标变量分组的单特征计数,用它做这个需求本身就不合适。 - 第二段catplot代码问题:
data = df[variable_list]只传入了特征列,没包含Survived,无法按目标分组Age是数值型,用catplot的count类型会把每个唯一值当成单独类别,导致柱子过度密集- 计数标签用
{0:.1f}K完全没必要,泰坦尼克数据集总样本仅几百,直接显示整数即可
三、快速探索简化方案
如果只追求快速出图、不需要精细美化,可以用pandas原生plot,代码更简洁:
for feat in features: # 分类特征/低基数数值特征用分组柱状图 if df[feat].nunique() < 10 or df[feat].dtype in ['object', 'category']: df.groupby(target)[feat].value_counts().unstack().plot(kind='bar', figsize=(8,4), title=f'{feat} by {target}') # 高基数数值特征用重叠直方图 else: df.groupby(target)[feat].plot(kind='hist', bins=15, alpha=0.5, legend=True, figsize=(8,4), title=f'{feat} by {target}') plt.show()
内容的提问来源于stack exchange,提问作者AbdullahQ
相关产品推荐
相关产品推荐

