基于双列条件过滤:移除特定条件下含NA值城市的所有行(dplyr实现)
Solution with dplyr
Got it, let's solve this problem using dplyr—here's a straightforward approach that fits your requirements perfectly:
First, make sure you have the dplyr package loaded:
library(dplyr)
Then, use this code to clean your data frame:
clean_df <- df %>% group_by(city) %>% # Filter out cities where ANY row with crime = "total" has an NA value filter(!any(crime == "total" & is.na(value))) %>% ungroup()
How this works:
group_by(city): Groups the data by each city, so we can evaluate conditions across all rows for a single city.filter(!any(crime == "total" & is.na(value))):- The
any()function checks if there's at least one row in the city's group wherecrimeis "total" andvalueis NA. - We use
!(NOT) to invert this check—so we keep only cities where this condition never occurs.
- The
ungroup(): Resets the grouping to return a standard tibble/data frame.
Result:
When you run this on your sample data, Amsterdam gets removed entirely (since its 2017 "total" crime value is NA), leaving only Rotterdam's rows:
clean_df #> # A tibble: 3 × 4 #> city year crime value #> <chr> <int> <chr> <int> #> 1 Rotterdam 2015 total 4901 #> 2 Rotterdam 2016 total 4830 #> 3 Rotterdam 2017 total 4659
This logic also works if your data has other crime types beyond "total"—any city with a NA in its "total" crime values will have all its rows removed, regardless of other crime categories.
内容的提问来源于stack exchange,提问作者Tdebeus
相关产品推荐
相关产品推荐

