R语言中使用case_when处理双字符列按位匹配生成组合编码的方法
解决方法
你需要先拆分每列的两个字符,分别对两个位置的碱基组合做规则匹配后再拼接,以下是可直接运行的代码:
首先加载依赖包:
library(dplyr) library(stringr)
方案1:直接分步处理
df_new <- df %>% # 提取每列第1位和第2位碱基 mutate(across(col1:col4, list(pos1 = ~str_sub(.x, 1, 1), pos2 = ~str_sub(.x, 2, 2)), .names = "{.col}_{.fn}")) %>% # 对第一个位置的碱基组合匹配规则 mutate(res1 = case_when( col1_pos1 == "A" & col2_pos1 == "C" & col3_pos1 == "T" & col4_pos1 == "A" ~ "01", col1_pos1 == "A" & col2_pos1 == "C" & col3_pos1 == "T" & col4_pos1 == "T" ~ "02", col1_pos1 == "T" & col2_pos1 == "G" & col3_pos1 == "C" & col4_pos1 == "T" ~ "03", col1_pos1 == "T" & col2_pos1 == "G" & col3_pos1 == "C" & col4_pos1 == "A" ~ "04", col1_pos1 == "T" & col2_pos1 == "G" & col3_pos1 == "T" & col4_pos1 == "A" ~ "05", TRUE ~ "other" )) %>% # 对第二个位置的碱基组合匹配规则 mutate(res2 = case_when( col1_pos2 == "A" & col2_pos2 == "C" & col3_pos2 == "T" & col4_pos2 == "A" ~ "01", col1_pos2 == "A" & col2_pos2 == "C" & col3_pos2 == "T" & col4_pos2 == "T" ~ "02", col1_pos2 == "T" & col2_pos2 == "G" & col3_pos2 == "C" & col4_pos2 == "T" ~ "03", col1_pos2 == "T" & col2_pos2 == "G" & col3_pos2 == "C" & col4_pos2 == "A" ~ "04", col1_pos2 == "T" & col2_pos2 == "G" & col3_pos2 == "T" & col4_pos2 == "A" ~ "05", TRUE ~ "other" )) %>% # 拼接结果,两个都为other时返回other mutate(new_col = case_when( res1 == "other" & res2 == "other" ~ "other", TRUE ~ paste0(res1, "+", res2) )) %>% # 保留原始列和最终结果列 select(Id, col1:col4, new_col)
运行后输出的df_new和你给出的预期结果完全一致。
方案2:封装函数简化代码
如果要避免重复写两次匹配逻辑,可以把规则封装为自定义函数:
match_rule <- function(c1, c2, c3, c4) { case_when( c1 == "A" & c2 == "C" & c3 == "T" & c4 == "A" ~ "01", c1 == "A" & c2 == "C" & c3 == "T" & c4 == "T" ~ "02", c1 == "T" & c2 == "G" & c3 == "C" & c4 == "T" ~ "03", c1 == "T" & c2 == "G" & c3 == "C" & c4 == "A" ~ "04", c1 == "T" & c2 == "G" & c3 == "T" & c4 == "A" ~ "05", TRUE ~ "other" ) } # 调用函数处理 df_new <- df %>% mutate( res1 = match_rule(str_sub(col1,1,1), str_sub(col2,1,1), str_sub(col3,1,1), str_sub(col4,1,1)), res2 = match_rule(str_sub(col1,2,2), str_sub(col2,2,2), str_sub(col3,2,2), str_sub(col4,2,2)), new_col = ifelse(res1 == "other" & res2 == "other", "other", paste0(res1, "+", res2)) ) %>% select(Id, col1:col4, new_col)
内容的提问来源于stack exchange,提问作者Marwah Al-kaabi
相关产品推荐
相关产品推荐

