在R语言中循环统计每行指定列的'yes'数量并生成新列
解决R语言数据框按规则填充新列的问题
需求说明
给数据框df新增status_check列,填充规则如下:
- 若
status列值为stable,则status_check设为stable - 若
status列值为Non-stable,统计该行所有以meds_开头的列中yes的数量:- 数量超过1 →
status_check设为Combo - 数量≤1 →
status_check设为Other
- 数量超过1 →
原始数据集
df <- data.frame(name=c("AJ", "DJ", "EJ", "MJ", "CJ"), meds_1=c("yes","yes", "no", "no", "yes"), meds_2=c("no", "no","no", "yes", "yes"), meds_3=c("no", "no","no", "no", "no"), meds_4=c("no", "no","no", "no", "no"), status=c("Non-stable","Non-stable","stable", "stable", "Non-stable")) # 初始化新列 df$status_check <- NA
解决方案
方法1:基础R向量化实现(推荐,效率更高)
无需循环,利用R的向量化特性直接处理整列:
# 筛选所有以meds_开头的列索引 meds_cols <- grep("^meds_", colnames(df)) # 按规则赋值 df$status_check <- ifelse( df$status == "stable", "stable", ifelse(rowSums(df[, meds_cols] == "yes") > 1, "Combo", "Other") )
方法2:tidyverse风格实现
如果习惯用tidyverse工具链,可借助dplyr的函数实现:
library(dplyr) df <- df %>% # 先统计每行meds_列的yes数量 mutate(meds_yes_count = rowSums(across(starts_with("meds_"), ~ .x == "yes"))) %>% # 按规则生成status_check mutate(status_check = case_when( status == "stable" ~ "stable", meds_yes_count > 1 ~ "Combo", TRUE ~ "Other" )) %>% # 可选:移除中间计数列 select(-meds_yes_count)
方法3:循环实现(不推荐大数据集)
如果一定要用循环,需修正原代码的索引问题,确保按行处理:
# 筛选meds_开头的列索引 meds_cols <- grep("^meds_", colnames(df)) for(i in 1:nrow(df)){ if(df$status[i] == "stable"){ df$status_check[i] <- "stable" } else { # 统计当前行meds_列的yes数量 yes_count <- sum(df[i, meds_cols] == "yes") df$status_check[i] <- ifelse(yes_count > 1, "Combo", "Other") } }
最终结果
执行上述任意方法后,输出结果如下:
print(df) # name meds_1 meds_2 meds_3 meds_4 status status_check # 1 AJ yes no no no Non-stable Other # 2 DJ yes no no no Non-stable Other # 3 EJ no no no no stable stable # 4 MJ no yes no no stable stable # 5 CJ yes yes no no Non-stable Combo
内容的提问来源于stack exchange,提问作者RStudent
相关产品推荐
相关产品推荐

