如何将以日龄为测量时间指标的长格式纵向数据转宽格式?
解决方案
要实现按首次/第二次测量转宽格式,核心是先给每个受试者的测量添加顺序标记,再进行转换,以下两种方法都可行:
方法1:基础R实现
- 先为每个受试者的测量添加顺序标签(1=首次,2=第二次):
dataset_col$time_point <- ave(dataset_col$agedays, dataset_col$subjid, FUN = function(x) order(x))
- 使用
reshape函数转换,此时timevar用新生成的顺序标签:
wide_data <- reshape(dataset_col, direction = "wide", idvar = "subjid", timevar = "time_point", v.names = c("cognitive_score", "motor_score"))
转换后的列名会是cognitive_score.1、motor_score.1、cognitive_score.2、motor_score.2,符合需求。
方法2:tidyverse工具实现(更直观易读)
如果习惯用tidyverse生态,步骤如下:
- 加载所需包:
library(dplyr) library(tidyr)
- 添加自定义顺序标签("first"=首次,"second"=第二次)并转宽:
wide_data <- dataset_col %>% group_by(subjid) %>% mutate(time_point = ifelse(agedays == min(agedays), "first", "second")) %>% ungroup() %>% pivot_wider( id_cols = subjid, names_from = time_point, values_from = c(cognitive_score, motor_score) )
转换后的列名会是cognitive_score_first、motor_score_first、cognitive_score_second、motor_score_second,可读性更强。
为什么原方法不符合需求?
你之前直接用agedays作为timevar,reshape函数会将agedays的具体数值作为列名后缀,而我们需要的是基于测量顺序的标记,所以必须先构造这个顺序变量再转换。
内容的提问来源于stack exchange,提问作者Makie
相关产品推荐
相关产品推荐

