You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr在另一列存在重复值时为某列创建新类别?

问题

我有一个记录各类研究坐标的dataframe,研究类型分为experiment(实验)和observation(观测),部分地点同时开展了这两种研究。我希望为这些地点的study列创建名为both的新类别,请问如何使用dplyr实现该需求?

示例数据

df1 <- data.frame(matrix(ncol = 4, nrow = 6))
colnames(df1)[1:4] <- c("value", "study", "lat","long")
df1$value <- c(1,1,2,3,4,4)
df1$study <- rep(c('experiment','observation'),3)
df1$lat <- c(37.541290,37.541290,38.936604,29.9511,51.509865,51.509865)
df1$long <- c(-77.434769,-77.434769,-119.986649,-90.0715,-0.118092,-0.118092)
df1

输出:

value       study      lat        long
1     1  experiment 37.54129  -77.434769
2     1 observation 37.54129  -77.434769
3     2  experiment 38.93660 -119.986649
4     3 observation 29.95110  -90.071500
5     4  experiment 51.50986   -0.118092
6     4 observation 51.50986   -0.118092

注:当study同时包含experiment和observation时,value列会出现重复值。

理想输出

value       study      lat        long
1     1        both 37.54129  -77.434769
2     2  experiment 38.93660 -119.986649
3     3 observation 29.95110  -90.071500
4     4        both 51.50986   -0.118092

解决方案

可以通过dplyr的分组、判断和去重操作实现,具体代码如下:

library(dplyr)

df_result <- df1 %>%
  # 按value、lat、long分组,标识同一地点
  group_by(value, lat, long) %>%
  # 判断分组内是否同时包含两种研究类型,设置对应study值
  mutate(study = ifelse(all(c("experiment", "observation") %in% study), 
                        "both", 
                        first(study))) %>%
  # 去重,保留每组唯一行
  distinct() %>%
  # 取消分组状态
  ungroup()

# 查看结果
df_result

代码说明

  • group_by(value, lat, long):将同一地点(相同value、经纬度)的行归为一组;
  • mutate(...):检查当前分组是否同时覆盖两种研究类型,是则将study设为both,否则保留原类型;
  • distinct():去除分组后重复的行,只保留每组一条记录;
  • ungroup():取消分组,恢复普通dataframe格式。

内容的提问来源于stack exchange,提问作者tassones

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 11:36:02