R语言按条件替换数据框值时组合代码失效问题咨询
问题原因及解决办法
为什么组合执行后全变成'HIGH'?
你应该是分两次执行了赋值操作,比如:
usdf$new_tests_per_thousand[usdf$new_tests_per_thousand < 3.22] <- 'LOW' usdf$new_tests_per_thousand[usdf$new_tests_per_thousand > 3.22] <- 'HIGH'
第一次替换后,这列的类型从数值型变成了字符型。第二次判断时,R会把数值3.22转成字符串'3.22',然后和列里的内容比较——字符比较是按ASCII码顺序来的,'LOW'的首字母L比'3'的ASCII码大,所以'LOW' > '3.22'会返回TRUE,结果所有非NA的行(包括已经改成'LOW'的)都被判定为大于3.22,最终全变成了'HIGH'。
正确的实现方式
方法1:用ifelse一次性完成判断
usdf$new_tests_per_thousand <- ifelse( usdf$new_tests_per_thousand < 3.22, 'LOW', ifelse(usdf$new_tests_per_thousand > 3.22, 'HIGH', NA) )
一次性完成所有条件判断,不会出现中间类型转换的问题,还能保留原有的NA值。
方法2:用dplyr::case_when(适合tidyverse用户)
library(dplyr) usdf <- usdf %>% mutate(new_tests_per_thousand = case_when( new_tests_per_thousand < 3.22 ~ 'LOW', new_tests_per_thousand > 3.22 ~ 'HIGH', TRUE ~ NA_character_ # 匹配剩下的情况(等于3.22或NA) ))
多条件场景下更清晰,逻辑一目了然。
方法3:分步操作但提前存索引
如果非要分步来,先把所有判断的索引提前算好,再统一赋值:
# 先计算符合条件的行索引 low_rows <- usdf$new_tests_per_thousand < 3.22 high_rows <- usdf$new_tests_per_thousand > 3.22 # 先把列转成字符型并设为NA usdf$new_tests_per_thousand <- NA_character_ # 再分别替换 usdf$new_tests_per_thousand[low_rows] <- 'LOW' usdf$new_tests_per_thousand[high_rows] <- 'HIGH'
这样就不会因为中间类型变化导致判断出错。
内容的提问来源于stack exchange,提问作者Donovan Morgan
相关产品推荐
相关产品推荐

