R语言数据清洗:mutate报错及病毒状态批量重编码求助
问题描述
我有一份包含不同地区物种病毒数据的大型数据集,示例数据如下:
Country ..2 Area Site ID Species Sample Original Sample/Specimen # <chr> <lgl> <chr> <chr> <chr> <chr> <chr> <chr> Tanzania NA UMNP UMNPhq AATPH PG Feces AATPHF2 Tanzania NA UMNP UMNPhq AATPI PG Feces AATPIF2 Tanzania NA UMNP UMNPhq AATPJ PG Feces AATPJF2 Tanzania NA UMNP UMNPhq ATTPK PG Feces ATTPKF2 Tanzania NA UMNP UMNPhq AATPL PG Feces AATPLF2 Filovirus (MOD) PCR Date (Filo MOD) <chr> <date> Indeterminant 2015-03-16 Indeterminant 2015-03-16 Indeterminant 2015-03-16 Indeterminant 2015-03-16 Negative 2015-03-16
数据集的dput输出:
structure(list(Country = c("Tanzania", "Tanzania", "Tanzania", "Tanzania", "Tanzania"), ...2 = c(NA, NA, NA, NA, NA), Area = c("UMNP", "UMNP", "UMNP", "UMNP", "UMNP"), Site = c("UMNPhq", "UMNPhq", "UMNPhq", "UMNPhq", "UMNPhq"), `Animal ID` = c("AATPH", "AATPI", "AATPJ", "ATTPK", "AATPL"), Species = c("Procolobus gordonorum", "Procolobus gordonorum", "Procolobus gordonorum", "Procolobus gordonorum", "Procolobus gordonorum"), `Sample Type` = c("Feces", "Feces", "Feces", "Feces", "Feces"), `Original Sample/Specimen #` = c("AATPHF2", "AATPIF2", "AATPJF2", "ATTPKF2", "AATPLF2"), `Filovirus (MOD) PCR` = c("Indeterminant", "Indeterminant", "Indeterminant", "Indeterminant", "Negative" ), `Date (Filo MOD)` = structure(c(16510, 16510, 16510, 16510, 16510), class = "Date")), row.names = c(NA, -5L), class = c("tbl_df", "tbl", "data.frame"))
我的需求:
- 筛选
Area为UMNP的数据 - 删除包含特定字符串的无关列(如
Performed by、Date of等) - 对病毒状态列(比如示例中的
Filovirus (MOD) PCR)批量重编码:将Indeterminant和Negative转为0,Positive转为1,且不能影响其他字符型列(如Country、Species等)
我尝试的代码运行后,所有非病毒状态的字符型列都变成了NA:
viral <- subset(data, Area %in% "UMNP") viralres <- viral %>% dplyr::select(-matches(c('Performed by ()', 'performed by', 'Date of', '1Performed by', 'Performed by', "Date ()", "...2"))) %>% mutate_if(is.character, ~case_when(. == "Indeterminant" ~ "0", . == "Negative" ~ "0", . == "Positive" ~ "1"))
解决方案
1. 问题根源
mutate_if(is.character, ...)会对所有字符型列执行重编码逻辑,而像Country、Species这类列的值既不是Indeterminant/Negative/Positive,在case_when中无匹配项,最终被转为NA。
2. 针对性解决方法
只对病毒状态列执行重编码,而非所有字符型列,以下是几种通用批量处理方式:
方法一:按列名特征匹配(推荐批量处理)
如果所有病毒状态列名称包含统一关键词(比如PCR),用matches()匹配目标列:
library(dplyr) viralres <- data %>% filter(Area == "UMNP") %>% # 统一用dplyr风格替代subset select(-matches(c('Performed by', 'Date of', 'Date \\(\\)', '\\.\\.\\.2'))) %>% # 转义正则特殊字符 mutate(across(matches("PCR"), ~case_when( . %in% c("Indeterminant", "Negative") ~ "0", . == "Positive" ~ "1", TRUE ~ . # 保留列中其他未匹配的状态值 )))
方法二:手动指定目标列
若病毒状态列无统一命名特征,直接列出目标列:
viralres <- data %>% filter(Area == "UMNP") %>% select(-matches(c('Performed by', 'Date of', 'Date \\(\\)', '\\.\\.\\.2'))) %>% mutate(across(c(`Filovirus (MOD) PCR`, "Other Virus PCR"), ~case_when( . %in% c("Indeterminant", "Negative") ~ "0", . == "Positive" ~ "1", TRUE ~ . )))
方法三:用recode简化逻辑
如果状态只有三种,用dplyr::recode让代码更简洁:
viralres <- data %>% filter(Area == "UMNP") %>% select(-matches(c('Performed by', 'Date of', 'Date \\(\\)', '\\.\\.\\.2'))) %>% mutate(across(matches("PCR"), ~recode(., "Indeterminant" = "0", "Negative" = "0", "Positive" = "1", .default = . # 保留其他未匹配值 )))
3. 关键说明
- 用
across()替代mutate_if(),精准控制操作列,避免误修改无关字符型列 - 正则表达式中,
()、.等特殊字符需用\\转义,否则会被当作正则语法解析 - 添加
TRUE ~ .或.default = .,可保留列中未匹配的其他状态值,避免转为NA
内容的提问来源于stack exchange,提问作者Marnee Roundtree
相关产品推荐
相关产品推荐

