按study分组重排data.frame失序数值变量 使其序号从1开始连续
R 按分组重编码为连续序号解决方案
核心逻辑
按指定分组列(本示例为study)拆分数据后,对需要处理的数值列,将其唯一值按升序排列后依次映射为从1开始的连续整数,其他列保持不变。
通用函数实现
1. tidyverse 版本(代码简洁易读)
依赖dplyr包,支持多分组列、多待重编码列:
library(dplyr) recode_seq_by_group <- function(df, group_col, recode_cols) { df %>% group_by(across(all_of(group_col))) %>% mutate(across(all_of(recode_cols), ~as.integer(factor(.x, levels = sort(unique(.x)))))) %>% ungroup() }
2. 基础R版本(无第三方依赖)
不需要安装任何包,可直接运行:
recode_seq_by_group_base <- function(df, group_col, recode_cols) { split_df <- split(df, df[[group_col]]) processed <- lapply(split_df, function(subdf) { for (col in recode_cols) { unique_vals <- sort(unique(subdf[[col]])) subdf[[col]] <- match(subdf[[col]], unique_vals) } return(subdf) }) do.call(rbind, c(processed, make.row.names = FALSE)) }
调用示例
# 加载示例数据 m=" study sample group outcome 1 1 1 A 1 1 1 B 1 1 2 A 1 1 2 B 1 3 1 A 1 3 1 B 1 3 2 A 1 3 2 B 2 1 2 A 2 1 2 B 2 2 2 A 2 2 2 B 2 3 2 A 2 3 2 B 3 1 1 A 3 1 1 B 3 1 2 A 3 1 2 B 3 2 1 A 3 2 1 B 3 2 2 A 3 2 2 B" data <- read.table(text=m, h=T) # 调用tidyverse版本函数处理 result <- recode_seq_by_group( df = data, group_col = "study", recode_cols = c("sample", "group") ) # 调用基础R版本函数处理 # result <- recode_seq_by_group_base(data, "study", c("sample", "group"))
结果验证
运行后得到的result与题目给出的期望输出完全一致,符合所有规则要求:
- study1的sample值3被重编码为2
- study2的group值2被重编码为1
- study3的所有值保持不变
- 其他列(outcome)无修改
内容的提问来源于stack exchange,提问作者Simon Harmel
相关产品推荐
相关产品推荐

