You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr的group_by和summarize按多列OR条件计算占比

解决方案

你可以通过以下代码实现需求:

首先定义目标数据框:

df <- tribble(
  ~id, ~x, ~y, ~z,
  'A',0,0,0,
  'A',1,0,0,
  'A',1,1,0,
  'B',0,0,0,
  'B',0,0,0,
  'B',1,0,0,
  'C',1,0,0,
  'C',0,0,0,
  'C',0,0,0,
  'C',1,0,0,
)

接下来用group_by()和summarize()计算每组的占比:

library(dplyr)

df %>%
  group_by(id) %>%
  summarise(
    result = round(mean(rowSums(across(c(x, y, z))) > 0) * 100, digits = 2)
  )

运行后输出结果:

# A tibble: 3 × 2
  id    result
  <chr>  <dbl>
1 A       66.7
2 B       33.3
3 C       50  

代码说明

  • across(c(x, y, z)):选中需要判断的x、y、z三列
  • rowSums(...) > 0:计算每行x/y/z的和,和大于0就代表该行至少有一列值为1,返回TRUE(对应1)或FALSE(对应0)
  • mean(...) * 100:利用逻辑值的数值特性,直接计算TRUE的占比并转为百分比
  • round(..., digits=2):将结果保留两位小数

内容的提问来源于stack exchange,提问作者rez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 06:04:54