如何结合循环与正则使用case_when?new_p变量赋值异常求助
我来帮你分析下问题的根源,然后给出几个针对不同场景的解决方案:
原代码的问题分析
你最初的循环写法有两个核心问题:
- 每次循环都会重新创建
newDT,直接覆盖上一次的结果,所以最后只保留了最后一次循环(对应lev[3] = "C")的赋值,其他行自然返回NA; case_when里只定义了当前lev[i]的匹配条件,没有处理其他类别,所以非当前类别的行都会被赋值为NA。
针对原数据集的解决方案
解法1:用dplyr批量生成case_when条件(易读性强)
不需要循环,直接一次性匹配所有d1的类别:
library(dplyr) library(purrr) tempDF <- structure(list(d1 = c("A", "B", "C"), d2 = c(40L, 50L, 20L), d3 = c(20L, 40L, 50L), d4 = c(60L, 30L, 30L), p_A = c(1L, 3L, 2L), p_B = c(3L, 4L, 3L), p_C = c(2L, 1L, 1L), p4 = c(5L, 5L, 4L)), class = "data.frame", row.names = c(NA, -3L)) lev <- levels(as.factor(tempDF$d1)) # 批量生成每个类别的匹配条件 case_conditions <- map(lev, ~ expr(d1 == !!.x ~ .data[[paste0("p_", !!.x)]])) %>% reduce(c) newDT <- tempDF %>% mutate(new_p = case_when(!!!case_conditions)) newDT
解释:用purrr::map为每个类别生成对应的case_when条件,再合并后传入case_when,一次性完成所有行的匹配。
解法2:矩阵索引(高效适合大数据集)
直接通过行号+目标列号的矩阵索引提取值,效率极高:
tempDF <- structure(list(d1 = c("A", "B", "C"), d2 = c(40L, 50L, 20L), d3 = c(20L, 40L, 50L), d4 = c(60L, 30L, 30L), p_A = c(1L, 3L, 2L), p_B = c(3L, 4L, 3L), p_C = c(2L, 1L, 1L), p4 = c(5L, 5L, 4L)), class = "data.frame", row.names = c(NA, -3L)) # 为每行生成对应的目标列名 target_cols <- paste0("p_", tempDF$d1) # 提取对应位置的值 tempDF$new_p <- tempDF[cbind(seq_len(nrow(tempDF)), match(target_cols, names(tempDF)))] tempDF
解释:先为每行生成匹配的p_*列名,再找到这些列在数据框中的位置,最后用矩阵索引直接提取对应值。
处理扩展数据集的函数错误
你后来用的函数出现警告和错误,原因是:
j <- match(paste0("p", "_", lev), names(tempDF))得到的是所有p_*列的位置(长度为3),而i <- match(tempDF$d1, lev)得到的是每行的类别索引(长度为5);- R会自动循环较短的向量,导致索引错位,最终取值错误。
修正后的高效解法(矩阵索引)
tempDF <- structure(list(d1 = c("A", "B", "C", "A", "C"), d2 = c(40L, 50L, 20L, 50L, 20L), d3 = c(20L, 40L, 50L, 40L, 50L), d4 = c(60L, 30L, 30L,60L, 30L), p_A = c(1L, 3L, 2L, 3L, 2L), p_B = c(3L, 4L, 3L, 3L, 4L), p_C = c(2L, 1L, 1L,2L, 1L), p4 = c(5L, 5L, 4L, 5L, 4L)), class = "data.frame", row.names = c(NA, -5L)) func <- function(tempDF){ # 为每行生成对应的目标列名 target_cols <- paste0("p_", tempDF$d1) # 找到每行目标列的位置 col_indices <- match(target_cols, names(tempDF)) # 矩阵索引提取值 tempDF$new_p <- tempDF[cbind(seq_len(nrow(tempDF)), col_indices)] return(tempDF) } newDT <- func(tempDF) newDT
易读性优先的tidyverse解法
用rowwise逐行处理,结合get函数提取对应列:
library(dplyr) tempDF <- structure(list(d1 = c("A", "B", "C", "A", "C"), d2 = c(40L, 50L, 20L, 50L, 20L), d3 = c(20L, 40L, 50L, 40L, 50L), d4 = c(60L, 30L, 30L,60L, 30L), p_A = c(1L, 3L, 2L, 3L, 2L), p_B = c(3L, 4L, 3L, 3L, 4L), p_C = c(2L, 1L, 1L,2L, 1L), p4 = c(5L, 5L, 4L, 5L, 4L)), class = "data.frame", row.names = c(NA, -5L)) newDT <- tempDF %>% rowwise() %>% mutate(new_p = get(paste0("p_", d1))) %>% ungroup() newDT
解释:rowwise()让mutate逐行执行,get()会根据当前行的d1值,提取对应的p_*列的数值,最后取消行分组恢复正常结构。
内容的提问来源于stack exchange,提问作者Krantz
相关产品推荐
相关产品推荐

