You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中将不同值的多行合并为单行的实现方法

R语言实现长格式转宽格式解决方案

方法一:使用tidyverse(tidyr + dplyr)

这是最常用的tidy风格处理方式,步骤清晰易读:

  1. 加载所需包:
library(tidyverse)
  1. 构造示例数据集(可替换为read.table/read.csv等导入你自己的数据):
df <- data.frame(
  ID = c("smith192", "green484", "green484", "tom584", "tom584", "tom584"),
  age = c(17, 12, 13, 12, 14, 15),
  disease = c("lung_cancer", "diabetes", "heart_irregularities", "colon_cancer", "diabetes", "malnutrition"),
  stringsAsFactors = FALSE
)
  1. 为每个ID的记录添加分组内序号,用来区分同ID下的不同观测:
df <- df %>%
  group_by(ID) %>%
  mutate(row_num = row_number()) %>%
  ungroup()
  1. 转换为宽格式,通过names_glue指定列名格式:
wide_df <- df %>%
  pivot_wider(
    id_cols = ID,
    names_from = row_num,
    values_from = c(age, disease),
    names_glue = "{.value}_{row_num}"
  )

运行后得到的wide_df就是目标格式,缺失位置会自动填充NA。

方法二:使用data.table(适合大数据集)

如果数据集规模较大,data.table的dcast处理效率更高:

  1. 加载data.table包:
library(data.table)
  1. 转换为data.table格式并添加分组序号:
dt <- as.data.table(df)
dt[, row_num := seq_len(.N), by = ID]
  1. 执行宽格式转换:
wide_dt <- dcast(
  dt,
  ID ~ row_num,
  value.var = c("age", "disease"),
  sep = "_"
)

输出结果与tidyverse方法一致,处理大样本时速度优势明显。

内容的提问来源于stack exchange,提问作者Aaron Z

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 15:45:31