You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何将数据框重复行并排合并并保留指定变量信息

R实现同组重复记录合并为单行方案

核心思路

按姓名、年龄、性别三个固定字段分组,为组内每条记录生成顺序编号,再将长表转换为宽表,按编号拆分体重和日期字段即可实现需求。

方法1:tidyverse 实现(推荐)

依赖dplyr和tidyr包,代码简洁易维护:

# 1. 构造示例原始数据(你可以替换成自己读取的数据)
df <- data.frame(
  id = c("01", "02", "03", "04", "05", "06"),
  name = c("ANA", "ANA", "KARLA", "KARLA", "MARIA", "MARIA"),
  age = c(18, 18, 21, 21, 30, 30),
  sex = rep("F", 6),
  weight = c(56.7, 59.5, 60.1, 60.4, 49.9, 50.2),
  date = as.Date(c("02/27/2020", "05/05/2020", "08/04/2020", "12/30/2020", "01/06/2020", "03/07/2020"), format = "%m/%d/%Y")
)

# 2. 加载依赖包
library(dplyr)
library(tidyr)

# 3. 数据处理
result <- df %>%
  # 按固定字段分组
  group_by(name, age, sex) %>%
  # 组内生成顺序编号,区分同组第1/2条记录
  mutate(rn = row_number()) %>%
  ungroup() %>%
  # 长表转宽表,拆分体重、日期字段
  pivot_wider(
    id_cols = c(name, age, sex),
    names_from = rn,
    values_from = c(weight, date)
  ) %>%
  # 生成新的连续id,放在最左侧
  mutate(id = row_number(), .before = 1) %>%
  # 调整列名匹配需求格式
  rename(
    weight1 = weight_1,
    date1 = date_1,
    weight2 = weight_2,
    date2 = date_2
  )

# 4. 可选:调整日期格式为月/日/年字符串格式
result$date1 <- format(result$date1, "%m/%d/%Y")
result$date2 <- format(result$date2, "%m/%d/%Y")

方法2:基础R实现(无需安装额外包)

# 1. 同上先构造/读取原始数据df

# 2. 按固定字段拆分数据
df_split <- split(df, list(df$name, df$age, df$sex), drop = TRUE)

# 3. 定义分组合并函数
combine_group <- function(x) {
  data.frame(
    id = NA,
    name = x$name[1],
    age = x$age[1],
    sex = x$sex[1],
    weight1 = x$weight[1],
    date1 = format(x$date[1], "%m/%d/%Y"),
    weight2 = x$weight[2],
    date2 = format(x$date[2], "%m/%d/%Y")
  )
}

# 4. 合并所有分组生成结果
result <- do.call(rbind, lapply(df_split, combine_group))
result$id <- seq(nrow(result))
rownames(result) <- NULL

以上两种方法均适配你给出的场景,若每组记录数不固定,只需调整pivot_wider的参数即可自动生成对应数量的体重、日期列。

内容的提问来源于stack exchange,提问作者Rita Carvalho Sauer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 18:27:04