R语言for循环条件判断含NA值报错,如何正确生成二分类变量?
问题原因
你的报错是因为在执行数值比较的if判断前没有提前排除NA值:当precipitation或temperature为NA时,df$precipitation[i] > 0这类比较运算的结果也是NA,if语句无法处理NA类型的逻辑值,就会抛出缺失值错误。你之前把NA判断放在了数值比较逻辑之后,因此NA还是会进入数值比较环节触发报错。
方案1:修改原for循环逻辑
调整判断顺序,先排查所有参与计算的变量是否存在NA,再走后续判断逻辑:
for(i in seq_len(nrow(df))){ # 先判断任意变量是否有NA,有则直接赋值NA if (is.na(df$snowheight[i]) | is.na(df$precipitation[i]) | is.na(df$temperature[i])) { df$dicht[i] <- NA } else if (df$snowheight[i] > 10) { # 满足snowheight>10才进后续判断 if (df$precipitation[i] <= 0) { df$dicht[i] <- 0 } else { df$dicht[i] <- ifelse(df$temperature[i] > 0, 1, 0) } } else { # snowheight<=10直接赋值NA df$dicht[i] <- NA } }
方案2:更高效的向量化实现(无需for循环)
用R的向量化运算直接批量生成结果,代码更简洁运行效率也更高:
基础R版本
df$dicht <- NA # 筛选无缺失且snowheight>10的行 valid_idx <- complete.cases(df[, c("snowheight", "precipitation", "temperature")]) & df$snowheight > 10 df[valid_idx, "dicht"] <- ifelse(df[valid_idx, "precipitation"] > 0 & df[valid_idx, "temperature"] > 0, 1, 0)
dplyr版本(更易读)
library(dplyr) df <- df %>% mutate(dicht = case_when( # 任意变量缺失或snowheight<=10时赋值NA is.na(snowheight) | is.na(precipitation) | is.na(temperature) | snowheight <= 10 ~ NA_real_, precipitation <= 0 ~ 0, precipitation > 0 & temperature > 0 ~ 1, TRUE ~ 0 ))
两种方案运行后得到的结果都和你给出的预期输出完全一致。
内容的提问来源于stack exchange,提问作者Zorin Ivanov
相关产品推荐
相关产品推荐

