You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr将重复观测长格式数据转换为多变量宽格式

长格式转多变量宽格式(R语言实现)

原始数据

ID      Approach Date
-42365 Sternotomy 18-11-2022
-42365 Thoracotomy 22-03-2024
-11234 Thoracotomy 12-03-2018
-11234 Sternotomy 17-05-2023

目标格式

ID     Approach_1    Date_1   Approach_2 Date_2
-42365 Sternotomy  18-11-2022 Thoracotomy 22-03-2024
-11234 Thoracotomy 12-03-2018 Sternotomy 17-05-2023

解决方案

方法1:dplyr + tidyr 组合

这是tidyverse生态下的标准做法,步骤清晰易读:

# 加载依赖包(若未安装先运行 install.packages(c("dplyr", "tidyr")))
library(dplyr)
library(tidyr)

# 构造示例数据(如果已有数据框可跳过此步)
df <- data.frame(
  ID = c(-42365, -42365, -11234, -11234),
  Approach = c("Sternotomy", "Thoracotomy", "Thoracotomy", "Sternotomy"),
  Date = c("18-11-2022", "22-03-2024", "12-03-2018", "17-05-2023")
)

# 分组添加序号,再转宽
df_wide <- df %>%
  group_by(ID) %>%
  mutate(row_id = row_number()) %>%  # 按原始顺序生成序号;若要按日期排序,先加 arrange(Date)
  ungroup() %>%
  pivot_wider(
    id_cols = ID,
    names_from = row_id,
    values_from = c(Approach, Date),
    names_glue = "{.value}_{row_id}"  # 自定义列名格式
  )

# 输出结果
df_wide

如果需要严格按日期先后生成_1/_2,在mutate前添加arrange(Date)即可。

方法2:data.table 实现

针对大数据集,data.table的处理效率更高:

# 加载包(未安装先运行 install.packages("data.table"))
library(data.table)

# 构造示例数据框并转为data.table
df <- data.table(
  ID = c(-42365, -42365, -11234, -11234),
  Approach = c("Sternotomy", "Thoracotomy", "Thoracotomy", "Sternotomy"),
  Date = c("18-11-2022", "22-03-2024", "12-03-2018", "17-05-2023")
)

# 添加分组序号并转宽
df_wide <- dcast(
  df[, row_id := seq_len(.N), by = ID],
  ID ~ row_id,
  value.var = c("Approach", "Date"),
  sep = "_"
)

# 调整列顺序以匹配目标格式
setcolorder(df_wide, c("ID", "Approach_1", "Date_1", "Approach_2", "Date_2"))

# 输出结果
df_wide

内容的提问来源于stack exchange,提问作者Rafik Margaryan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 07:44:59