R语言中dplyr::filter()与base::subset()筛选误删过多观测值问题
Hey there! Let's break down why you're seeing more rows get deleted than expected when negating your two-condition filter in dplyr::filter() or base R's subset().
The Root Cause: Missing Values (NA)
Here's the key issue: when you run your positive filter (e.g., filter(df, cond1 & cond2)), both filter() and subset() only keep rows where the condition evaluates to TRUE. They automatically drop rows where the condition is FALSE or NA.
But when you negate that condition (e.g., filter(df, !(cond1 & cond2))), rows where the original condition was NA will still evaluate to NA after negation. These rows get dropped too—on top of the 92 rows you intended to exclude. That's why you're losing more data than expected.
How to Fix It
You have two reliable ways to get the exact set of rows you want (all rows except the 92 matching your positive conditions):
1. Explicitly Include NA Rows in Your Negated Condition
Adjust your filter to explicitly keep rows where either the negated condition is TRUE, or any of the original condition variables have NA values. For example, if your conditions are var1 == "X" and var2 > 50:
# Using dplyr df %>% filter(!(var1 == "X" & var2 > 50) | is.na(var1) | is.na(var2)) # Using base R subset(df, !(var1 == "X" & var2 > 50) | is.na(var1) | is.na(var2))
2. Use Row Indices to Exclude Only the Matching Rows
A more straightforward approach is to first identify the rows that match your positive condition, then exclude only those rows from the original dataset. This preserves all other rows (including those with NA values):
# dplyr method matching_rows <- df %>% filter(var1 == "X" & var2 > 50) %>% pull(row_number()) result <- df %>% slice(-matching_rows) # Base R method matching_rows <- which(with(df, var1 == "X" & var2 > 50)) result <- df[-matching_rows, ]
Verify the Issue
To confirm NA values are the culprit, check how many rows have NA in your filter variables:
# Count rows with NA in either condition variable sum(is.na(df$var1) | is.na(df$var2))
This number should match the extra rows you're losing (total rows deleted minus 92).
内容的提问来源于stack exchange,提问作者Wokkel

