You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于条件为多行数据批量分配分组值(R语言)

解决方案

方法1:使用dplyr(推荐,代码简洁高效)

首先加载dplyr包(未安装需先运行install.packages("dplyr")):

library(dplyr)

按规则修正分组:

# 基于首次评估年龄修正分组
myinput_fixed <- myinput %>%
  group_by(ID) %>%
  arrange(timepoint) %>% # 按评估年份排序,确保取到首次评估数据
  mutate(
    first_age = first(age),
    group = ifelse(first_age < 50, "young", "old"),
    timepoint = row_number() # 可选:将每个ID的评估时间转为顺序编号
  ) %>%
  ungroup() %>%
  select(-first_age) # 移除辅助列

# 查看修正结果
print(myinput_fixed)

输出结果:

# A tibble: 7 × 4
  ID        timepoint   age group 
  <chr>         <int> <dbl> <chr> 
1 person1          1    49 young 
2 person1          2    50 young 
3 person2          1    45 young 
4 person3          1    48 young 
5 person3          2    49 young 
6 person3          3    50 young 
7 person4          1    55 old   

方法2:基础R实现(无需额外包)

如果不想依赖dplyr,可使用基础R代码:

# 计算每个ID的首次评估年龄
first_age <- sapply(unique(myinput$ID), function(id) {
  id_rows <- myinput[myinput$ID == id, ]
  id_rows$age[which.min(id_rows$timepoint)]
})

# 创建分组映射表
group_map <- ifelse(first_age < 50, "young", "old")

# 修正原数据的分组列
myinput$group <- group_map[myinput$ID]

# 可选:将每个ID的评估时间转为顺序编号
myinput$timepoint <- ave(myinput$timepoint, myinput$ID, FUN = function(x) rank(x, ties.method = "first"))

# 查看结果
print(myinput)

逻辑说明

核心规则是参与者分组由首次评估时的年龄永久决定,后续年龄变化不影响分组。两种方法都能高效处理2万行规模的数据,dplyr版本更易读和维护。

内容的提问来源于stack exchange,提问作者Nereus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 18:16:33