如何按列名筛选:移除名称含_1且全为NA/空值的列?
移除R数据框中满足特定条件的列
原始数据框定义
a <- c(3, 2, 1) a_1 <- c(NA, "", NA) b <- c(3, 4, 1) b_1 <- c(3, NA, 4) c <- c("", "", "") c_1 <- c(5, 8, 9) d <- c(6, 9, 10) d_1 <- c("", "", "") e <- c(NA, NA, NA) e_1 <- c(NA, NA, NA) df <- data.frame(a, a_1, b, b_1, c, c_1, d, d_1, e, e_1)
需求
需要移除列名包含"_1"且所有元素均为空值或NA的列,保留其他所有列(包括非"_1"结尾但全为空/NA的列,比如c、e)。
错误的处理方式
之前使用的代码会移除所有全为空/NA的列,不符合需求:
empty_columns <- colSums(is.na(df) | df == "") == nrow(df) df[, !empty_columns] df <- df[, colSums(is.na(df)) < nrow(df)]
运行后结果:
a b b_1 c_1 d 1 3 3 3 5 6 2 2 4 NA 8 9 3 1 1 4 9 10
正确解决方案
需要同时筛选出满足「列名含"_1"」和「列内全为空/NA」的列,仅移除这些列:
# 判断每列是否全为空值或NA is_all_empty <- colSums(is.na(df) | df == "") == nrow(df) # 判断列名是否包含"_1" has_underscore1 <- grepl("_1", colnames(df)) # 确定需要移除的列:同时满足上述两个条件 cols_to_remove <- is_all_empty & has_underscore1 # 保留不需要移除的列 df_cleaned <- df[, !cols_to_remove]
期望结果
运行上述代码后,输出的df_cleaned如下:
a b b_1 c c_1 d e 1 3 3 3 5 6 NA 2 2 4 NA 8 9 NA 3 1 1 4 9 10 NA
内容的提问来源于stack exchange,提问作者hy9fesh
相关产品推荐
相关产品推荐

