You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr计算各大学诺贝尔奖不同奖项类型的占比分布?

计算各大学诺贝尔奖奖项类型占比的实现方法

核心思路

先统计每个大学各奖项类型的获奖次数,再基于大学总获奖数计算对应占比,最终得到每个大学内部不同奖项的占比分布。

代码实现

library(dplyr)

# 1. 统计每个大学各奖项的获奖次数,保留大学分组
uni_category_counts <- nobel %>%
  group_by(name_of_university, category) %>%
  summarise(count = n(), .groups = "drop_last")

# 2. 计算总获奖数与各奖项占比
uni_category_percentages <- uni_category_counts %>%
  mutate(
    total_awards = sum(count),
    percentage = (count / total_awards) * 100
  ) %>%
  ungroup()

# 可选:将百分比格式化为带%的字符串(如75.0%)
uni_category_percentages <- uni_category_percentages %>%
  mutate(percentage_formatted = sprintf("%.1f%%", percentage))

代码解释

  • group_by(name_of_university, category):同时按大学和奖项类型分组,确保能统计每个大学单个奖项的获奖次数
  • summarise(count = n(), .groups = "drop_last"):统计每组的获奖次数,仅取消最内层的category分组,保留大学分组用于后续总次数计算
  • total_awards = sum(count):在大学分组内计算该大学的总获奖数
  • percentage = (count / total_awards) * 100:计算当前奖项类型在该大学总奖项中的占比
  • sprintf("%.1f%%", percentage):将数值型百分比格式化为带百分号的字符串,提升可读性

示例输出

以你提到的鸭子大学为例,运行代码后会得到如下结果(仅展示关键列):

name_of_universitycategorycounttotal_awardspercentagepercentage_formatted
鸭子大学文学3475.075.0%
鸭子大学物理1425.025.0%

内容的提问来源于stack exchange,提问作者Julio1000MX

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 00:20:38