如何基于Allele与Match的唯一配对筛选R数据框子集?
Hey there! Let's figure out why your code isn't working and get you the right result for filtering rows where the Allele + Match pair only appears once.
First, let's formalize your data into a proper data frame (I assume you might have skipped this step in your snippet):
Value <- c(FALSE, TRUE, FALSE, TRUE, FALSE) Allele <- c('a','a','a','b','b') Match <- c('b','b','c','b','b') x <- data.frame(Value, Allele, Match)
What's wrong with your original code?
The duplicated() function operates on data objects (like columns of a data frame), but you passed it two string literals: "Allele" and "Match". This treats those two strings as a short vector with no duplicates, so duplicated() returns c(FALSE, FALSE). When you use ! on that and subset your data frame, R recycles the TRUE values to match all rows—hence you get the full original data back.
Fixes that work
Option 1: Base R with duplicated() (check both directions)
To check for duplicate pairs of Allele and Match, you need to pass those two columns as a single object to duplicated(). Also, since duplicated() only marks the second and later occurrences of a duplicate, we need to check from the end too to catch the first occurrence of a repeated pair:
# Identify all rows where the Allele-Match pair is duplicated (in either direction) duplicate_rows <- duplicated(x[, c("Allele", "Match")]) | duplicated(x[, c("Allele", "Match")], fromLast = TRUE) # Keep only rows that are NOT part of a duplicated pair filtered_x <- x[!duplicate_rows, ]
Running this gives you the only row where the pair is unique:
Value Allele Match 3 FALSE a c
(Quick note: You mentioned expecting FALSE,a,b, but in your data, the a-b pair appears twice (rows 1 and 2), so it's actually not unique. The truly unique pair is a-c in row 3!)
Option 2: Tidyverse/dplyr approach (more readable)
If you prefer using the tidyverse, grouping by the pair and filtering for groups with only one row is super intuitive:
library(dplyr) filtered_x <- x %>% group_by(Allele, Match) %>% filter(n() == 1) %>% ungroup()
This does the exact same thing as the base R method, and the code reads like plain English: group by the two columns, keep groups where the number of rows (n()) is 1, then remove the grouping.
内容的提问来源于stack exchange,提问作者Ascaris

