如何在R中统计多值分组字段中各变量的重复次数?
Hey there! The problem you're facing is that your sub column contains comma-separated combined values, so group_by(sub) treats each full string as a distinct group—hence why "A" and "A,B" get counted separately. Let's fix this so you get the individual counts you're after (A:3, B:3, C:2)!
Solution with dplyr + tidyr
First, we'll use separate_rows() from the tidyr package to split those comma-separated values into individual rows, then count each unique value as usual.
- First, let's replicate your dataset for testing:
mydata <- data.frame( num = 1:6, sub = c("A", "A,B", "C", "A", "B", "B,C") )
- Now run the processing code:
library(dplyr) library(tidyr) # Split comma-separated values into individual rows, then count mydata2 <- mydata %>% separate_rows(sub, sep = ",") %>% group_by(sub) %>% summarise(n = n()) %>% ungroup() # Optional: removes grouping structure if you don't need it # View the result mydata2
This will give you exactly what you want:
# A tibble: 3 × 2 sub n <chr> <int> 1 A 3 2 B 3 3 C 2
Base R Alternative
If you prefer not to load extra packages, you can use base R functions to achieve the same result:
# Split each string into a list, then flatten to a vector sub_values <- unlist(strsplit(mydata$sub, ",")) # Count occurrences with table() count_table <- table(sub_values) # Convert to data frame if you want a structure similar to the dplyr output as.data.frame(count_table)
This will return:
sub_values Freq 1 A 3 2 B 3 3 C 2
内容的提问来源于stack exchange,提问作者Jay Choi
相关产品推荐
相关产品推荐

