You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ggplot:如何筛选特定阈值以上数据并展示正确百分比

解决方案

1. 数据预处理

核心目标是先计算分组下4/5分的占比,再基于分组变量绘制子图,示例默认使用Python的pandas+seaborn+matplotlib栈实现。

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt

# 生成高喜爱度标记列:得分4、5记为True,其余为False
df['high_liking'] = df['Liking of cats'].isin([4, 5])

# 按分类维度(示例为性别gender、是否有车has_car)计算占比,输出百分比格式
ratio_df = df.groupby(['gender', 'has_car'], as_index=False)['high_liking'].mean()
ratio_df['high_liking'] = ratio_df['high_liking'].round(4) * 100

2. 子图绘制

方案1:单分类变量分面展示

适合多分类变量独立展示的场景,使用seaborn的FacetGrid自动生成分面子图:

# 示例:按性别分面展示高喜爱度占比
g = sns.FacetGrid(df, col='gender', height=4, aspect=0.9)
g.map_dataframe(
    sns.barplot,
    y='high_liking',
    estimator=lambda x: sum(x)/len(x)*100,
    errorbar=None,
    palette='Set2'
)
g.set_axis_labels('', '喜爱度得分≥4的占比(%)')
g.set_titles('性别:{col_name}')
plt.show()

方案2:多分类变量交叉子图

适合同时展示多个分类维度的对比,手动控制子图布局:

# 示例:1行2列分别展示性别、是否有车两个维度的占比
fig, axes = plt.subplots(1, 2, figsize=(10, 4), sharey=True)

# 第一个子图:按性别分组
sns.barplot(
    data=ratio_df, x='gender', y='high_liking',
    ax=axes[0], palette='Set2'
)
axes[0].set_title('按性别分组')
axes[0].set_ylabel('喜爱度得分≥4的占比(%)')
axes[0].set_xlabel('性别')

# 第二个子图:按是否有车分组
sns.barplot(
    data=ratio_df, x='has_car', y='high_liking',
    ax=axes[1], palette='Set2'
)
axes[1].set_title('按是否拥有汽车分组')
axes[1].set_xlabel('是否拥有汽车')

plt.tight_layout()
plt.show()

可选调整项

  • 如果需要单独展示4分、5分各自的占比而非合并,可以将high_liking拆分为独立标记列,或绘图时添加hue='Liking of cats'参数区分两个得分
  • 子图行列数可以根据分类变量数量调整plt.subplots()的前两个参数,超过2个分类变量可改用行+列的网格布局
  • 百分比精度可以通过round()函数的参数自行调整

内容的提问来源于stack exchange,提问作者ag123987

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 23:57:02