如何通过循环按col1分组统计col2唯一值数量并生成col3列?
按分组计算唯一值数量并新增列的解决方案
首先构造与你提供一致的示例数据:
data <- data.frame( col1 = c(1,1,1,2,2,3), col2 = c(5,5,3,10,11,11) )
方法一:循环实现
如果你想用循环遍历所有分组,可按以下步骤操作:
# 获取col1的所有唯一分组值 unique_groups <- unique(data$col1) # 初始化新增列col3 data$col3 <- NA # 遍历每个分组,计算并赋值 for (group in unique_groups) { # 计算当前分组col2的唯一值数量 count <- length(unique(data$col2[data$col1 == group])) # 给当前分组的所有行设置col3值 data$col3[data$col1 == group] <- count }
运行后得到的结果:
| col1 | col2 | col3 |
|---|---|---|
| 1 | 5 | 2 |
| 1 | 5 | 2 |
| 1 | 3 | 2 |
| 2 | 10 | 2 |
| 2 | 11 | 2 |
| 3 | 11 | 1 |
方法二:用dplyr实现(更简洁高效)
在R中处理分组操作,推荐使用dplyr包的函数,代码更简洁且性能更优:
# 首次使用需安装dplyr # install.packages("dplyr") library(dplyr) data <- data %>% group_by(col1) %>% mutate(col3 = n_distinct(col2)) %>% # n_distinct直接计算唯一值数量 ungroup()
这个方法无需手动写循环,分组后直接给每组新增对应的唯一值数量列,结果和循环方法完全一致。
内容的提问来源于stack exchange,提问作者Sverdo
相关产品推荐
相关产品推荐

