如何按组筛选包含指定两个字符串值的R语言数据集?
Hey there! Let's figure out why your code isn't working and fix it right away.
问题根源
Your current filter(grepl("^PVDA$&^GL$", var)) checks if individual rows have a var value that's both exactly "PVDA" and exactly "GL" at the same time— which is impossible, since each row's var can only hold one value. That's why no rows are being kept!
Instead, we need to check if an entire city group contains both "PVDA" and "GL" somewhere in its var values, then keep all rows from those groups.
正确解决方案
Here are two straightforward ways to do this with tidyverse tools:
方法1:使用%in%和all()(适合检查多个值的场景)
This is clean and scalable if you ever need to check for more than two values later:
library(tidyverse) # 示例数据集 df <- tribble( ~city, ~var, "A", "PVDA", "A", "GL", "A", "GMBL", "B", "GL", "B", "VVD", "C", "CDA", "C", "VVD" ) # 筛选出同时包含PVDA和GL的城市组 df %>% group_by(city) %>% filter(all(c("PVDA", "GL") %in% var)) %>% ungroup() # 可选,取消分组如果后续不需要
方法2:使用两个any()(更直观,适合少量值的检查)
If you prefer something more explicit for two values, this works too:
df %>% group_by(city) %>% filter(any(var == "PVDA") & any(var == "GL")) %>% ungroup()
输出结果
Both methods will return exactly what you want—all rows for city A:
# A tibble: 3 × 2 city var <chr> <chr> 1 A PVDA 2 A GL 3 A GMBL
内容的提问来源于stack exchange,提问作者Tdebeus

