ggplot:如何筛选特定阈值以上数据并展示正确百分比
解决方案
1. 数据预处理
核心目标是先计算分组下4/5分的占比,再基于分组变量绘制子图,示例默认使用Python的pandas+seaborn+matplotlib栈实现。
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt # 生成高喜爱度标记列:得分4、5记为True,其余为False df['high_liking'] = df['Liking of cats'].isin([4, 5]) # 按分类维度(示例为性别gender、是否有车has_car)计算占比,输出百分比格式 ratio_df = df.groupby(['gender', 'has_car'], as_index=False)['high_liking'].mean() ratio_df['high_liking'] = ratio_df['high_liking'].round(4) * 100
2. 子图绘制
方案1:单分类变量分面展示
适合多分类变量独立展示的场景,使用seaborn的FacetGrid自动生成分面子图:
# 示例:按性别分面展示高喜爱度占比 g = sns.FacetGrid(df, col='gender', height=4, aspect=0.9) g.map_dataframe( sns.barplot, y='high_liking', estimator=lambda x: sum(x)/len(x)*100, errorbar=None, palette='Set2' ) g.set_axis_labels('', '喜爱度得分≥4的占比(%)') g.set_titles('性别:{col_name}') plt.show()
方案2:多分类变量交叉子图
适合同时展示多个分类维度的对比,手动控制子图布局:
# 示例:1行2列分别展示性别、是否有车两个维度的占比 fig, axes = plt.subplots(1, 2, figsize=(10, 4), sharey=True) # 第一个子图:按性别分组 sns.barplot( data=ratio_df, x='gender', y='high_liking', ax=axes[0], palette='Set2' ) axes[0].set_title('按性别分组') axes[0].set_ylabel('喜爱度得分≥4的占比(%)') axes[0].set_xlabel('性别') # 第二个子图:按是否有车分组 sns.barplot( data=ratio_df, x='has_car', y='high_liking', ax=axes[1], palette='Set2' ) axes[1].set_title('按是否拥有汽车分组') axes[1].set_xlabel('是否拥有汽车') plt.tight_layout() plt.show()
可选调整项
- 如果需要单独展示4分、5分各自的占比而非合并,可以将
high_liking拆分为独立标记列,或绘图时添加hue='Liking of cats'参数区分两个得分 - 子图行列数可以根据分类变量数量调整
plt.subplots()的前两个参数,超过2个分类变量可改用行+列的网格布局 - 百分比精度可以通过
round()函数的参数自行调整
内容的提问来源于stack exchange,提问作者ag123987
相关产品推荐
相关产品推荐

