如何基于另一列唯一值筛选列值?R语言数据集处理求助
Hey there! Let's sort out this filtering problem for you. The root issue with your current code is that all(x == x) will always return TRUE—every value is equal to itself, so it doesn't actually check if all values in a Class group are identical. That's why you're getting all rows back instead of just the groups you want.
Here are two straightforward solutions to get the result you need:
1. Base R Approach
We can adjust the ave() function to properly check if all values in each Class group are the same, using either of these methods:
Option A: Check unique value count
# First, identify which classes have only one unique Value valid_classes <- unique(dataset$Class[ave(dataset$Value, dataset$Class, FUN = function(x) length(unique(x)) == 1)]) # Filter the dataset to keep only those valid classes sub <- dataset[dataset$Class %in% valid_classes, ]
Option B: Compare all values to the first in the group
This is a more concise way to run the check directly in the row index:
sub <- dataset[ave(dataset$Value, dataset$Class, FUN = function(x) all(x == x[1])), ]
2. dplyr Approach (More Readable)
If you use the dplyr package, the logic becomes super intuitive—we group by Class, then filter to keep only groups where there's exactly one distinct Value:
library(dplyr) sub <- dataset %>% group_by(Class) %>% filter(n_distinct(Value) == 1) %>% ungroup()
Expected Result
Running either solution will return only rows for Class A, C, and D (all their Value entries are identical), and exclude Class B (which has two different Value entries):
| Class | Value |
|---|---|
| A | 5.4 |
| A | 5.4 |
| A | 5.4 |
| C | 4.02 |
| C | 4.02 |
| C | 4.02 |
| D | 6.33 |
| D | 6.33 |
内容的提问来源于stack exchange,提问作者Adam Amin

