R语言:如何排除指定变量组全为NA的数据行?
Hey there! Let's sort out that error you ran into when trying to filter your dataframe. The problem with your original code is two-fold: missing a critical comma in the row index, and checking for the wrong condition (you want to keep rows that aren't all NA in columns 4-22, not the ones that are).
Base R Solution
Here's the corrected base R code that does exactly what you need:
# Keep rows where columns 4-22 are NOT all NA df <- df[rowSums(is.na(df[, 4:22])) != 19, ]
Let me break this down:
df[, 4:22]targets the columns you care about (4 through 22, which is 19 total columns)is.na()converts all NA values in those columns toTRUErowSums()counts how many NAs are in each row of those columns- We keep rows where that count isn't 19 (meaning at least one value is non-NA)
- The trailing comma
[, ]ensures we keep all columns (including your first 3, regardless of their NA status)
Tidyverse (dplyr) Solution
If you prefer a more readable approach using the tidyverse, this works too—no need to calculate the number of columns:
library(dplyr) df <- df %>% filter(if_any(4:22, ~ !is.na(.x)))
This code directly checks: "Does this row have at least one non-NA value in columns 4-22?" If yes, it stays. Super straightforward!
Either of these methods will keep your first 3 columns intact, ignore their NA status, and only remove rows where columns 4-22 are completely filled with NAs.
内容的提问来源于stack exchange,提问作者J. Doe

