You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言数据清洗:mutate报错及病毒状态批量重编码求助

问题描述

我有一份包含不同地区物种病毒数据的大型数据集,示例数据如下:

Country    ..2  Area    Site    ID      Species Sample    Original Sample/Specimen #
<chr>     <lgl> <chr>   <chr>   <chr>   <chr>   <chr>    <chr>
Tanzania    NA  UMNP    UMNPhq  AATPH   PG     Feces    AATPHF2 
Tanzania    NA  UMNP    UMNPhq  AATPI   PG     Feces    AATPIF2 
Tanzania    NA  UMNP    UMNPhq  AATPJ   PG     Feces    AATPJF2 
Tanzania    NA  UMNP    UMNPhq  ATTPK   PG     Feces    ATTPKF2 
Tanzania    NA  UMNP    UMNPhq  AATPL   PG     Feces    AATPLF2 

Filovirus (MOD) PCR  Date (Filo MOD)
<chr>                <date>
Indeterminant        2015-03-16
Indeterminant        2015-03-16
Indeterminant        2015-03-16
Indeterminant        2015-03-16
Negative             2015-03-16

数据集的dput输出:

structure(list(Country = c("Tanzania", "Tanzania", "Tanzania", 
"Tanzania", "Tanzania"), ...2 = c(NA, NA, NA, NA, NA), Area = c("UMNP", 
"UMNP", "UMNP", "UMNP", "UMNP"), Site = c("UMNPhq", "UMNPhq", 
"UMNPhq", "UMNPhq", "UMNPhq"), `Animal ID` = c("AATPH", "AATPI", 
"AATPJ", "ATTPK", "AATPL"), Species = c("Procolobus gordonorum", 
"Procolobus gordonorum", "Procolobus gordonorum", "Procolobus gordonorum", 
"Procolobus gordonorum"), `Sample Type` = c("Feces", "Feces", 
"Feces", "Feces", "Feces"), `Original Sample/Specimen #` = c("AATPHF2", 
"AATPIF2", "AATPJF2", "ATTPKF2", "AATPLF2"), `Filovirus (MOD) PCR` = c("Indeterminant", 
"Indeterminant", "Indeterminant", "Indeterminant", "Negative"
), `Date (Filo MOD)` = structure(c(16510, 16510, 16510, 16510, 
16510), class = "Date")), row.names = c(NA, -5L), class = c("tbl_df", 
"tbl", "data.frame"))

我的需求:

  • 筛选Area为UMNP的数据
  • 删除包含特定字符串的无关列(如Performed by、Date of等)
  • 对病毒状态列(比如示例中的Filovirus (MOD) PCR)批量重编码:将Indeterminant和Negative转为0,Positive转为1,且不能影响其他字符型列(如Country、Species等)

我尝试的代码运行后,所有非病毒状态的字符型列都变成了NA:

viral <- subset(data, Area %in% "UMNP")

viralres <- viral %>% 
     dplyr::select(-matches(c('Performed by ()', 'performed by', 'Date of', '1Performed by', 'Performed by', "Date ()", "...2"))) %>%
    mutate_if(is.character, ~case_when(. == "Indeterminant" ~ "0", 
                                       . == "Negative" ~ "0", 
                                       . == "Positive" ~ "1"))
解决方案

1. 问题根源

mutate_if(is.character, ...)会对所有字符型列执行重编码逻辑,而像Country、Species这类列的值既不是Indeterminant/Negative/Positive,在case_when中无匹配项,最终被转为NA。

2. 针对性解决方法

只对病毒状态列执行重编码,而非所有字符型列,以下是几种通用批量处理方式:

方法一:按列名特征匹配(推荐批量处理)

如果所有病毒状态列名称包含统一关键词(比如PCR),用matches()匹配目标列:

library(dplyr)

viralres <- data %>%
  filter(Area == "UMNP") %>% # 统一用dplyr风格替代subset
  select(-matches(c('Performed by', 'Date of', 'Date \\(\\)', '\\.\\.\\.2'))) %>% # 转义正则特殊字符
  mutate(across(matches("PCR"), ~case_when(
    . %in% c("Indeterminant", "Negative") ~ "0",
    . == "Positive" ~ "1",
    TRUE ~ . # 保留列中其他未匹配的状态值
  )))

方法二:手动指定目标列

若病毒状态列无统一命名特征,直接列出目标列:

viralres <- data %>%
  filter(Area == "UMNP") %>%
  select(-matches(c('Performed by', 'Date of', 'Date \\(\\)', '\\.\\.\\.2'))) %>%
  mutate(across(c(`Filovirus (MOD) PCR`, "Other Virus PCR"), ~case_when(
    . %in% c("Indeterminant", "Negative") ~ "0",
    . == "Positive" ~ "1",
    TRUE ~ .
  )))

方法三:用recode简化逻辑

如果状态只有三种,用dplyr::recode让代码更简洁:

viralres <- data %>%
  filter(Area == "UMNP") %>%
  select(-matches(c('Performed by', 'Date of', 'Date \\(\\)', '\\.\\.\\.2'))) %>%
  mutate(across(matches("PCR"), ~recode(.,
    "Indeterminant" = "0",
    "Negative" = "0",
    "Positive" = "1",
    .default = . # 保留其他未匹配值
  )))

3. 关键说明

  • 用across()替代mutate_if(),精准控制操作列,避免误修改无关字符型列
  • 正则表达式中,()、.等特殊字符需用\\转义,否则会被当作正则语法解析
  • 添加TRUE ~ .或.default = .,可保留列中未匹配的其他状态值,避免转为NA

内容的提问来源于stack exchange,提问作者Marnee Roundtree

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 19:25:28