基于名称含指定字符串的多列筛选R数据框
Hey there! Totally get where you're coming from—dealing with variable numbers of columns that follow a naming pattern is such a common pain point when you're getting comfortable with R. Let's break down some simple, scalable solutions for your problem.
First, let's recap your sample data so everyone's on the same page:
data <- data.frame( name=c("aaa","bbb","ccc","ddd"), 'type_01'=c("match", NA, NA, "match"), 'type_02'=c("part",NA,"match","match"), 'type_03'=c(NA,NA,NA,"part") )
方法1:Base R 实现
You're already on the right track using grep() to target your "type" columns. We can build on that to avoid hardcoding each column name:
- 先定位所有含"type"的列:
type_cols <- grep("type", names(data))
- 筛选所有"type"列全为NA的行:
这里有两种直观的方式:
- 用
rowSums()统计每行NA的数量,等于列数就说明全是NA:
# 筛选结果 filtered_data <- data[rowSums(is.na(data[, type_cols])) == length(type_cols), ]
- 用
apply()逐行检查是否所有值都是NA:
filtered_data <- data[apply(is.na(data[, type_cols]), 1, all), ]
方法2:Tidyverse (dplyr) 实现
If you're using the tidyverse, dplyr has a super clean way to handle this with if_all() and contains()—no need to mess with column indices:
library(dplyr) filtered_data <- data %>% filter(if_all(contains("type"), is.na))
扩展到其他条件
Great news—both methods work for any condition, not just checking NA! For example, if you wanted to keep rows where all "type" columns are "match":
- Base R:
filtered_data <- data[rowSums(data[, type_cols] == "match") == length(type_cols), ]
- Tidyverse:
filtered_data <- data %>% filter(if_all(contains("type"), ~ .x == "match"))
All these solutions will automatically adapt no matter how many "type" columns you have (whether it's 3 or 20)—no more manually listing each column!
内容的提问来源于stack exchange,提问作者Nerdbert

