R语言for循环筛选data.frame时报错missing value where TRUE/FALSE needed
问题解答
报错原因
该报错的核心原因是你的studenti数据集的sesso列存在NA(缺失值):当执行if (studenti$sesso[i] == "M")判断时,若当前行sesso为NA,NA == "M"的返回结果也是NA,而非TRUE/FALSE布尔值,if语句无法识别NA值就会触发该错误。你补充的dput输出仅为数据集前5行,这部分数据没有缺失值,因此问题出在你完整数据集的后续行中。
解决方案
你可以根据需求选择以下任意一种方案实现筛选sesso为M的行的需求:
方案1:修改现有循环逻辑
仅需在判断条件中优先排除缺失值即可,注意循环内标量判断使用&&逻辑运算符:
maschi = data.frame() indice = 0 for (i in seq_along(studenti$sesso)) { if (!is.na(studenti$sesso[i]) && studenti$sesso[i] == "M") { indice = indice + 1 maschi[indice,] <- studenti[i,] } }
方案2:使用R原生向量化筛选(更推荐)
R的向量化操作比循环效率更高,一行代码即可完成筛选,向量判断使用&逻辑运算符:
maschi <- studenti[!is.na(studenti$sesso) & studenti$sesso == "M", ]
方案3:使用dplyr包筛选(语法更简洁)
如果使用tidyverse生态的dplyr包,filter函数会自动忽略判断结果为NA的行,无需额外处理缺失值:
library(dplyr) maschi <- studenti %>% filter(sesso == "M")
内容的提问来源于stack exchange,提问作者Gennaro Arguzzi
相关产品推荐
相关产品推荐

