You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:按指定公式计算数据框中分类因子的频率占比

R语言实现方案

你给出的公式存在括号缺失的笔误,正确计算逻辑应为每个ID分组下,(Y和O的Counts之和 ÷ 该ID所有Counts总和) × 100,以下是两种可直接运行的实现方法:

方法1:使用dplyr包(语法简洁易读)

# 未安装包先运行 install.packages("dplyr")
library(dplyr)

df <- data.frame( ID = c("A","A","A","B","B","B","C","C","C"), 
                  levels = c( "Y", "R", "O","Y", "R", "O","Y", "R", "O" ),
                  Counts=c(5,1,5,10,2,1,3,5,8))

result <- df %>%
  group_by(ID) %>%
  summarise(
    total = sum(Counts),
    sum_YO = sum(Counts[levels %in% c("Y", "O")]),
    freq = sum_YO / total * 100
  ) %>%
  select(ID, freq)

运行后得到的实际计算结果如下:

ID  freq
  <chr> <dbl>
1 A      90.9
2 B      84.6
3 C      68.8

你给出的预期输出示例数值为占位值,可根据实际业务需求调整公式细节。

方法2:基础R实现(无需安装额外依赖)

df <- data.frame( ID = c("A","A","A","B","B","B","C","C","C"), 
                  levels = c( "Y", "R", "O","Y", "R", "O","Y", "R", "O" ),
                  Counts=c(5,1,5,10,2,1,3,5,8))

result <- do.call(rbind, lapply(split(df, df$ID), function(group_df) {
  total_count <- sum(group_df$Counts)
  yo_count <- sum(group_df$Counts[group_df$levels %in% c("Y", "O")])
  data.frame(ID = unique(group_df$ID), freq = yo_count / total_count * 100)
}))
rownames(result) <- NULL

内容的提问来源于stack exchange,提问作者Marwah Al-kaabi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 00:24:05