You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中按分组计算计数占总数的百分比

问题描述

我想在R语言中计算数据框各因子水平的计数占总计数的百分比,同时保留分组结构,但一直没成功。

我能拿到总计数:

df %>% summarise(sum(num))
# 15

也能拿到分组后的各水平计数:

df %>% group_by(species) %>% summarise(sum(num))
# A tibble: 3 × 2
#   species                  `sum(num)`
#   <chr>                         <int>
# 1 Farfantepenaeus duorarum          4
# 2 Farfantepenaeus notialis          0
# 3 Farfantepenaeus spp              11

但想要得到如下格式的输出却做不到:

#   species                     Percent
#   <chr>                         <chr>
# 1 Farfantepenaeus duorarum       4 / 15 = 0.267
# 2 Farfantepenaeus notialis       0 / 15 = 0.000
# 3 Farfantepenaeus spp           11 / 15 = 0.733

我最接近的尝试是用reframe(),但它返回的是无分组的数据,还丢了species列:

df %>% group_by(species) %>% 
  summarise(factor_count=sum(num)) %>% 
  # 提示:请用`reframe()`替代,但`reframe()`总是返回无分组数据
  reframe(percent=factor_count/sum(df$num))

# A tibble: 3 × 1
  percent
    <dbl>
1   0.267
2   0    
3   0.733

数据如下:

> dput(df)
structure(list(species = c("Farfantepenaeus notialis", "Farfantepenaeus spp", 
"Farfantepenaeus notialis", "Farfantepenaeus notialis", "Farfantepenaeus duorarum", 
"Farfantepenaeus duorarum", "Farfantepenaeus notialis", "Farfantepenaeus spp", 
"Farfantepenaeus duorarum", "Farfantepenaeus spp", "Farfantepenaeus notialis", 
"Farfantepenaeus duorarum", "Farfantepenaeus spp", "Farfantepenaeus notialis", 
"Farfantepenaeus notialis", "Farfantepenaeus spp", "Farfantepenaeus duorarum", 
"Farfantepenaeus spp", "Farfantepenaeus spp", "Farfantepenaeus duorarum", 
"Farfantepenaeus duorarum", "Farfantepenaeus spp", "Farfantepenaeus spp", 
"Farfantepenaeus spp", "Farfantepenaeus notialis"), num = c(0L, 
0L, 0L, 0L, 1L, 0L, 0L, 2L, 0L, 3L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 
0L, 0L, 0L, 3L, 0L, 2L, 4L, 0L)), row.names = c(159897L, 174698L, 
236857L, 190237L, 327321L, 272931L, 304567L, 75538L, 109206L, 
351373L, 280332L, 163966L, 282183L, 341197L, 316962L, 354703L, 
343971L, 95333L, 244258L, 254061L, 87561L, 186908L, 221318L, 
258688L, 97737L), class = "data.frame")

解决方案

你不需要用reframe(),直接在summarise()里完成所有计算即可,这里提供两种可行方法:

方法1:提前定义总计数,生成格式化百分比

先计算全局总计数,再在分组汇总时生成指定格式的百分比列:

library(dplyr)

# 先计算总计数
total_count <- sum(df$num)

df %>%
  group_by(species) %>%
  summarise(
    Percent = sprintf("%d / %d = %.3f", sum(num), total_count, sum(num)/total_count)
  )

输出结果:

# A tibble: 3 × 2
  species                  Percent          
  <chr>                    <chr>            
1 Farfantepenaeus duorarum 4 / 15 = 0.267    
2 Farfantepenaeus notialis 0 / 15 = 0.000    
3 Farfantepenaeus spp     11 / 15 = 0.733

方法2:无需提前定义变量,汇总后计算总计数

如果不想单独定义总计数变量,可以在分组汇总后,用mutate()计算总计数并生成百分比:

df %>%
  group_by(species) %>%
  summarise(factor_count = sum(num)) %>%
  mutate(
    total = sum(factor_count),
    Percent = sprintf("%d / %d = %.3f", factor_count, total, factor_count/total)
  ) %>%
  select(species, Percent)

此方法会先得到分组计数,再通过sum(factor_count)获取总计数,最后生成目标格式的百分比列,再筛选出需要的列输出。

为什么之前的reframe()方法失效?

reframe()的设计逻辑就是返回无分组的数据框,且你只指定了percent列,所以它会丢弃原有的species列。实际上用summarise()就能完成所有需求,不需要切换到reframe()。


内容的提问来源于stack exchange,提问作者Nate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 14:05:59