基于其他列1的出现次数重新赋值score列的实现方案咨询
问题描述
给定如下R数据集:
structure(list(ID = c(1, 2, 3, 4, 6, 7), V = c(0, 0, 1, 1, 1, 0), Mus = c(1, 0, 1, 1, 1, 0), R = c(1, 0, 1, 1, 1, 1), E = c(1, 0, 0, 1, 0, 0), S = c(1, 0, 1, 1, 1, 0), t = c(0, 0, 0, 1, 0, 0), score = c(1, 0.4, 1, 0.4, 0.4, 0.4)), row.names = c(NA, -6L), class = c("tbl_df", "tbl", "data.frame"), na.action = structure(c(`5` = 5L, `12` = 12L, `15` = 15L, `21` = 21L, `22` = 22L, `23` = 23L, `34` = 34L, `44` = 44L, `46` = 46L, `52` = 52L, `56` = 56L, `57` = 57L, `58` = 58L ), class = "omit"))
需按以下规则重新赋值score列:
- 若当前行除
ID和score外的列中,数字1的出现次数>3,score设为1; - 若次数=3,
score设为0.4; - 若次数<3,
score设为0。
方法1:for循环实现
先将数据存入变量,再逐行计算1的个数并赋值:
# 存储数据集 df <- structure(list(ID = c(1, 2, 3, 4, 6, 7), V = c(0, 0, 1, 1, 1, 0), Mus = c(1, 0, 1, 1, 1, 0), R = c(1, 0, 1, 1, 1, 1), E = c(1, 0, 0, 1, 0, 0), S = c(1, 0, 1, 1, 1, 0), t = c(0, 0, 0, 1, 0, 0), score = c(1, 0.4, 1, 0.4, 0.4, 0.4)), row.names = c(NA, -6L), class = c("tbl_df", "tbl", "data.frame"), na.action = structure(c(`5` = 5L, `12` = 12L, `15` = 15L, `21` = 21L, `22` = 22L, `23` = 23L, `34` = 34L, `44` = 44L, `46` = 46L, `52` = 52L, `56` = 56L, `57` = 57L, `58` = 58L), class = "omit")) # 循环处理每一行 for (i in 1:nrow(df)) { # 统计目标列中1的数量 count_ones <- sum(df[i, c("V", "Mus", "R", "E", "S", "t")] == 1) # 按规则赋值 if (count_ones > 3) { df$score[i] <- 1 } else if (count_ones == 3) { df$score[i] <- 0.4 } else { df$score[i] <- 0 } }
方法2:dplyr实现
用rowwise()做行操作,搭配case_when()写条件逻辑,代码更简洁:
library(dplyr) df <- df %>% rowwise() %>% mutate( count_ones = sum(c_across(V:t) == 1), # 统计V到t列的1的数量 score = case_when( count_ones > 3 ~ 1, count_ones == 3 ~ 0.4, TRUE ~ 0 ) ) %>% ungroup() # 取消行分组 # 若不需要保留count_ones列,可追加: # select(-count_ones)
方法3:apply函数实现
用apply()按行处理目标列,直接返回赋值结果:
# 选取需要统计的列(排除ID和score) target_cols <- df[, c("V", "Mus", "R", "E", "S", "t")] # 按行计算并赋值 df$score <- apply(target_cols, 1, function(x) { cnt <- sum(x == 1) if (cnt > 3) 1 else if (cnt == 3) 0.4 else 0 })
方法4:purrr::map实现
用map_dbl()遍历每行数据,计算后返回数值型结果:
library(purrr) # 方法4.1:拆分数据为行列表后处理 target_cols <- df[, c("V", "Mus", "R", "E", "S", "t")] row_list <- split(target_cols, seq(nrow(target_cols))) df$score <- map_dbl(row_list, function(row) { cnt <- sum(row == 1) if (cnt > 3) 1 else if (cnt == 3) 0.4 else 0 }) # 方法4.2:结合dplyr用pmap_dbl df <- df %>% mutate( score = pmap_dbl(select(., V:t), function(...) { cnt <- sum(c(...) == 1) if (cnt > 3) 1 else if (cnt == 3) 0.4 else 0 }) )
内容的提问来源于stack exchange,提问作者12666727b9
相关产品推荐
相关产品推荐

